open-swe/ui/src/routeTree.gen.ts

544 lines
18 KiB
TypeScript
Raw Normal View History

feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
/* eslint-disable */
// @ts-nocheck
// noinspection JSUnusedGlobalSymbols
// This file was automatically generated by TanStack Router.
// You should NOT make any changes in this file as it will be overwritten.
// Additionally, you should also exclude this file from your linter and/or formatter to prevent it from being checked or modified.
import { Route as rootRouteImport } from './routes/__root'
import { Route as UsageRouteImport } from './routes/usage'
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
import { Route as ReviewRouteImport } from './routes/review'
import { Route as MySettingsRouteImport } from './routes/my-settings'
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
import { Route as LoginRouteImport } from './routes/login'
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
import { Route as IntegrationsRouteImport } from './routes/integrations'
import { Route as CloudAgentsRouteImport } from './routes/cloud-agents'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
import { Route as AgentsRouteImport } from './routes/agents'
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
import { Route as AdminRouteImport } from './routes/admin'
import { Route as IndexRouteImport } from './routes/index'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
import { Route as AgentsIndexRouteImport } from './routes/agents/index'
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
import { Route as ReviewStylesRouteImport } from './routes/review_.styles'
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
import { Route as AgentsSnapshotsRouteImport } from './routes/agents_.snapshots'
import { Route as AgentsInstructionsRouteImport } from './routes/agents_.instructions'
import { Route as AgentsThreadsRouteImport } from './routes/agents/threads'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
import { Route as AgentsThreadIdRouteImport } from './routes/agents/$threadId'
import { Route as AdminEvalsRouteImport } from './routes/admin_.evals'
import { Route as AgentsReviewsIndexRouteImport } from './routes/agents/reviews/index'
import { Route as AgentsAutomationsIndexRouteImport } from './routes/agents/automations/index'
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
import { Route as ReviewRepositoriesOwnerRouteImport } from './routes/review_.repositories.$owner'
import { Route as AgentsAutomationsNewRouteImport } from './routes/agents/automations/new'
import { Route as AgentsAutomationsScheduleIdRouteImport } from './routes/agents/automations/$scheduleId'
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
import { Route as AgentsThreadIdPlanRouteImport } from './routes/agents/$threadId_.plan'
import { Route as AgentsReviewsOwnerRepoNumberRouteImport } from './routes/agents/reviews/$owner.$repo.$number'
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
const UsageRoute = UsageRouteImport.update({
id: '/usage',
path: '/usage',
getParentRoute: () => rootRouteImport,
} as any)
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
const ReviewRoute = ReviewRouteImport.update({
id: '/review',
path: '/review',
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) * feat: tune reviewer for precision — web/wiki tools + recalibrated prompt Reviewer agent now has web_search, fetch_url, and http_request alongside the finding tools, so it can verify library semantics and consult the DeepWiki auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>) before flagging cross-file or architectural concerns. Prompt rewritten to push precision over recall: - explicit severity ladder pushing reviews toward bimodal high/low instead of defaulting to medium - ≤200-char description target (gold set averages ~186 chars; we were at ~436) - mandatory docs / wiki / code lookup before flagging concurrency, security, or perf — the three categories that dominated false positives - "do not flag" list covering compiler/linter-catchable nits, speculative claims without a concrete attacker/interleaving/scale, style preferences the codebase doesn't share, and test-quality nits on non-test diffs - smart file-selection guidance for large PRs (deprioritize generated / vendored / pure-rename hunks) Eval config switched to openai:gpt-5.5 + high reasoning effort for the next benchmark run. * trim prompt * subagent prompting * confidence ratings * added medium * enforce confidence threshold * . * reviewer: precision-tuned prompt + drop confidence gate Rewrites the reviewer system prompt around a defensibility bar (anchor + failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file list (style nits, speculation, scope-policing, same-bug fan-out), and a checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26% style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new prompt targets each class directly. Confidence is still recorded on every finding for post-hoc calibration but no longer gates publication — the audit showed the gate was a no-op (agent self-rated 65% of findings "high" regardless), and the prompt's defensibility bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD, the confidence_threshold kwarg on filter_findings_for_publish, the confidence_filtered score_mode, and the min_confidence kwarg on the eval target's _extract_comments — all dead once the gate is gone. Also removes the "informational" severity tier from the Severity enum, SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for FYI observations the dataset never rewards. * benchmax * adding google provider * slight steering * tuning * more tuning * fix * cleanup * reducing overfitting * Add per-repo review style profiles and inject them into the reviewer. Dashboard users can analyze historical PR review feedback per repository, edit the resulting style guide, and have it loaded from LangGraph Store at reviewer runtime (including Martian eval runs) keyed by owner/name. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix review style job errors leaking exception details to clients. Return generic dashboard messages while logging full stack traces server-side. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 11:35:00 -07:00
getParentRoute: () => rootRouteImport,
} as any)
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
const MySettingsRoute = MySettingsRouteImport.update({
id: '/my-settings',
path: '/my-settings',
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
getParentRoute: () => rootRouteImport,
} as any)
const LoginRoute = LoginRouteImport.update({
id: '/login',
path: '/login',
getParentRoute: () => rootRouteImport,
} as any)
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
const IntegrationsRoute = IntegrationsRouteImport.update({
id: '/integrations',
path: '/integrations',
getParentRoute: () => rootRouteImport,
} as any)
const CloudAgentsRoute = CloudAgentsRouteImport.update({
id: '/cloud-agents',
path: '/cloud-agents',
getParentRoute: () => rootRouteImport,
} as any)
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
const AgentsRoute = AgentsRouteImport.update({
id: '/agents',
path: '/agents',
getParentRoute: () => rootRouteImport,
} as any)
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
const AdminRoute = AdminRouteImport.update({
id: '/admin',
path: '/admin',
getParentRoute: () => rootRouteImport,
} as any)
const IndexRoute = IndexRouteImport.update({
id: '/',
path: '/',
getParentRoute: () => rootRouteImport,
} as any)
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
const AgentsIndexRoute = AgentsIndexRouteImport.update({
id: '/',
path: '/',
getParentRoute: () => AgentsRoute,
} as any)
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
const ReviewStylesRoute = ReviewStylesRouteImport.update({
id: '/review_/styles',
path: '/review/styles',
getParentRoute: () => rootRouteImport,
} as any)
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
const AgentsSnapshotsRoute = AgentsSnapshotsRouteImport.update({
id: '/agents_/snapshots',
path: '/agents/snapshots',
getParentRoute: () => rootRouteImport,
} as any)
const AgentsInstructionsRoute = AgentsInstructionsRouteImport.update({
id: '/agents_/instructions',
path: '/agents/instructions',
getParentRoute: () => rootRouteImport,
} as any)
const AgentsThreadsRoute = AgentsThreadsRouteImport.update({
id: '/threads',
path: '/threads',
getParentRoute: () => AgentsRoute,
} as any)
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
const AgentsThreadIdRoute = AgentsThreadIdRouteImport.update({
id: '/$threadId',
path: '/$threadId',
getParentRoute: () => AgentsRoute,
} as any)
const AdminEvalsRoute = AdminEvalsRouteImport.update({
id: '/admin_/evals',
path: '/admin/evals',
getParentRoute: () => rootRouteImport,
} as any)
const AgentsReviewsIndexRoute = AgentsReviewsIndexRouteImport.update({
id: '/reviews/',
path: '/reviews/',
getParentRoute: () => AgentsRoute,
} as any)
const AgentsAutomationsIndexRoute = AgentsAutomationsIndexRouteImport.update({
id: '/automations/',
path: '/automations/',
getParentRoute: () => AgentsRoute,
} as any)
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
const ReviewRepositoriesOwnerRoute = ReviewRepositoriesOwnerRouteImport.update({
id: '/review_/repositories/$owner',
path: '/review/repositories/$owner',
getParentRoute: () => rootRouteImport,
} as any)
const AgentsAutomationsNewRoute = AgentsAutomationsNewRouteImport.update({
id: '/automations/new',
path: '/automations/new',
getParentRoute: () => AgentsRoute,
} as any)
const AgentsAutomationsScheduleIdRoute =
AgentsAutomationsScheduleIdRouteImport.update({
id: '/automations/$scheduleId',
path: '/automations/$scheduleId',
getParentRoute: () => AgentsRoute,
} as any)
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
const AgentsThreadIdPlanRoute = AgentsThreadIdPlanRouteImport.update({
id: '/$threadId_/plan',
path: '/$threadId/plan',
getParentRoute: () => AgentsRoute,
} as any)
const AgentsReviewsOwnerRepoNumberRoute =
AgentsReviewsOwnerRepoNumberRouteImport.update({
id: '/reviews/$owner/$repo/$number',
path: '/reviews/$owner/$repo/$number',
getParentRoute: () => AgentsRoute,
} as any)
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
export interface FileRoutesByFullPath {
'/': typeof IndexRoute
'/admin': typeof AdminRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents': typeof AgentsRouteWithChildren
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
'/cloud-agents': typeof CloudAgentsRoute
'/integrations': typeof IntegrationsRoute
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
'/login': typeof LoginRoute
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
'/my-settings': typeof MySettingsRoute
'/review': typeof ReviewRoute
'/usage': typeof UsageRoute
'/admin/evals': typeof AdminEvalsRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents/$threadId': typeof AgentsThreadIdRoute
'/agents/threads': typeof AgentsThreadsRoute
'/agents/instructions': typeof AgentsInstructionsRoute
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
'/agents/snapshots': typeof AgentsSnapshotsRoute
'/review/styles': typeof ReviewStylesRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents/': typeof AgentsIndexRoute
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
'/agents/$threadId/plan': typeof AgentsThreadIdPlanRoute
'/agents/automations/$scheduleId': typeof AgentsAutomationsScheduleIdRoute
'/agents/automations/new': typeof AgentsAutomationsNewRoute
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
'/review/repositories/$owner': typeof ReviewRepositoriesOwnerRoute
'/agents/automations/': typeof AgentsAutomationsIndexRoute
'/agents/reviews/': typeof AgentsReviewsIndexRoute
'/agents/reviews/$owner/$repo/$number': typeof AgentsReviewsOwnerRepoNumberRoute
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
}
export interface FileRoutesByTo {
'/': typeof IndexRoute
'/admin': typeof AdminRoute
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
'/cloud-agents': typeof CloudAgentsRoute
'/integrations': typeof IntegrationsRoute
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
'/login': typeof LoginRoute
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
'/my-settings': typeof MySettingsRoute
'/review': typeof ReviewRoute
'/usage': typeof UsageRoute
'/admin/evals': typeof AdminEvalsRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents/$threadId': typeof AgentsThreadIdRoute
'/agents/threads': typeof AgentsThreadsRoute
'/agents/instructions': typeof AgentsInstructionsRoute
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
'/agents/snapshots': typeof AgentsSnapshotsRoute
'/review/styles': typeof ReviewStylesRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents': typeof AgentsIndexRoute
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
'/agents/$threadId/plan': typeof AgentsThreadIdPlanRoute
'/agents/automations/$scheduleId': typeof AgentsAutomationsScheduleIdRoute
'/agents/automations/new': typeof AgentsAutomationsNewRoute
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
'/review/repositories/$owner': typeof ReviewRepositoriesOwnerRoute
'/agents/automations': typeof AgentsAutomationsIndexRoute
'/agents/reviews': typeof AgentsReviewsIndexRoute
'/agents/reviews/$owner/$repo/$number': typeof AgentsReviewsOwnerRepoNumberRoute
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
}
export interface FileRoutesById {
__root__: typeof rootRouteImport
'/': typeof IndexRoute
'/admin': typeof AdminRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents': typeof AgentsRouteWithChildren
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
'/cloud-agents': typeof CloudAgentsRoute
'/integrations': typeof IntegrationsRoute
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
'/login': typeof LoginRoute
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
'/my-settings': typeof MySettingsRoute
'/review': typeof ReviewRoute
'/usage': typeof UsageRoute
'/admin_/evals': typeof AdminEvalsRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents/$threadId': typeof AgentsThreadIdRoute
'/agents/threads': typeof AgentsThreadsRoute
'/agents_/instructions': typeof AgentsInstructionsRoute
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
'/agents_/snapshots': typeof AgentsSnapshotsRoute
'/review_/styles': typeof ReviewStylesRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents/': typeof AgentsIndexRoute
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
'/agents/$threadId_/plan': typeof AgentsThreadIdPlanRoute
'/agents/automations/$scheduleId': typeof AgentsAutomationsScheduleIdRoute
'/agents/automations/new': typeof AgentsAutomationsNewRoute
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
'/review_/repositories/$owner': typeof ReviewRepositoriesOwnerRoute
'/agents/automations/': typeof AgentsAutomationsIndexRoute
'/agents/reviews/': typeof AgentsReviewsIndexRoute
'/agents/reviews/$owner/$repo/$number': typeof AgentsReviewsOwnerRepoNumberRoute
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
}
export interface FileRouteTypes {
fileRoutesByFullPath: FileRoutesByFullPath
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
fullPaths:
| '/'
| '/admin'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
| '/agents'
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
| '/cloud-agents'
| '/integrations'
| '/login'
| '/my-settings'
| '/review'
| '/usage'
| '/admin/evals'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
| '/agents/$threadId'
| '/agents/threads'
| '/agents/instructions'
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
| '/agents/snapshots'
| '/review/styles'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
| '/agents/'
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
| '/agents/$threadId/plan'
| '/agents/automations/$scheduleId'
| '/agents/automations/new'
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
| '/review/repositories/$owner'
| '/agents/automations/'
| '/agents/reviews/'
| '/agents/reviews/$owner/$repo/$number'
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
fileRoutesByTo: FileRoutesByTo
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
to:
| '/'
| '/admin'
| '/cloud-agents'
| '/integrations'
| '/login'
| '/my-settings'
| '/review'
| '/usage'
| '/admin/evals'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
| '/agents/$threadId'
| '/agents/threads'
| '/agents/instructions'
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
| '/agents/snapshots'
| '/review/styles'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
| '/agents'
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
| '/agents/$threadId/plan'
| '/agents/automations/$scheduleId'
| '/agents/automations/new'
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
| '/review/repositories/$owner'
| '/agents/automations'
| '/agents/reviews'
| '/agents/reviews/$owner/$repo/$number'
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
id:
| '__root__'
| '/'
| '/admin'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
| '/agents'
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
| '/cloud-agents'
| '/integrations'
| '/login'
| '/my-settings'
| '/review'
| '/usage'
| '/admin_/evals'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
| '/agents/$threadId'
| '/agents/threads'
| '/agents_/instructions'
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
| '/agents_/snapshots'
| '/review_/styles'
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
| '/agents/'
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
| '/agents/$threadId_/plan'
| '/agents/automations/$scheduleId'
| '/agents/automations/new'
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
| '/review_/repositories/$owner'
| '/agents/automations/'
| '/agents/reviews/'
| '/agents/reviews/$owner/$repo/$number'
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
fileRoutesById: FileRoutesById
}
export interface RootRouteChildren {
IndexRoute: typeof IndexRoute
AdminRoute: typeof AdminRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
AgentsRoute: typeof AgentsRouteWithChildren
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
CloudAgentsRoute: typeof CloudAgentsRoute
IntegrationsRoute: typeof IntegrationsRoute
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
LoginRoute: typeof LoginRoute
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
MySettingsRoute: typeof MySettingsRoute
ReviewRoute: typeof ReviewRoute
UsageRoute: typeof UsageRoute
AdminEvalsRoute: typeof AdminEvalsRoute
AgentsInstructionsRoute: typeof AgentsInstructionsRoute
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
AgentsSnapshotsRoute: typeof AgentsSnapshotsRoute
ReviewStylesRoute: typeof ReviewStylesRoute
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
ReviewRepositoriesOwnerRoute: typeof ReviewRepositoriesOwnerRoute
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
}
declare module '@tanstack/react-router' {
interface FileRoutesByPath {
'/usage': {
id: '/usage'
path: '/usage'
fullPath: '/usage'
preLoaderRoute: typeof UsageRouteImport
parentRoute: typeof rootRouteImport
}
'/review': {
id: '/review'
path: '/review'
fullPath: '/review'
preLoaderRoute: typeof ReviewRouteImport
parentRoute: typeof rootRouteImport
}
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
'/my-settings': {
id: '/my-settings'
path: '/my-settings'
fullPath: '/my-settings'
preLoaderRoute: typeof MySettingsRouteImport
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
parentRoute: typeof rootRouteImport
}
'/login': {
id: '/login'
path: '/login'
fullPath: '/login'
preLoaderRoute: typeof LoginRouteImport
parentRoute: typeof rootRouteImport
}
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
'/integrations': {
id: '/integrations'
path: '/integrations'
fullPath: '/integrations'
preLoaderRoute: typeof IntegrationsRouteImport
parentRoute: typeof rootRouteImport
}
'/cloud-agents': {
id: '/cloud-agents'
path: '/cloud-agents'
fullPath: '/cloud-agents'
preLoaderRoute: typeof CloudAgentsRouteImport
parentRoute: typeof rootRouteImport
}
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents': {
id: '/agents'
path: '/agents'
fullPath: '/agents'
preLoaderRoute: typeof AgentsRouteImport
parentRoute: typeof rootRouteImport
}
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
'/admin': {
id: '/admin'
path: '/admin'
fullPath: '/admin'
preLoaderRoute: typeof AdminRouteImport
parentRoute: typeof rootRouteImport
}
'/': {
id: '/'
path: '/'
fullPath: '/'
preLoaderRoute: typeof IndexRouteImport
parentRoute: typeof rootRouteImport
}
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents/': {
id: '/agents/'
path: '/'
fullPath: '/agents/'
preLoaderRoute: typeof AgentsIndexRouteImport
parentRoute: typeof AgentsRoute
}
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
'/review_/styles': {
id: '/review_/styles'
path: '/review/styles'
fullPath: '/review/styles'
preLoaderRoute: typeof ReviewStylesRouteImport
parentRoute: typeof rootRouteImport
}
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
'/agents_/snapshots': {
id: '/agents_/snapshots'
path: '/agents/snapshots'
fullPath: '/agents/snapshots'
preLoaderRoute: typeof AgentsSnapshotsRouteImport
parentRoute: typeof rootRouteImport
}
'/agents_/instructions': {
id: '/agents_/instructions'
path: '/agents/instructions'
fullPath: '/agents/instructions'
preLoaderRoute: typeof AgentsInstructionsRouteImport
parentRoute: typeof rootRouteImport
}
'/agents/threads': {
id: '/agents/threads'
path: '/threads'
fullPath: '/agents/threads'
preLoaderRoute: typeof AgentsThreadsRouteImport
parentRoute: typeof AgentsRoute
}
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
'/agents/$threadId': {
id: '/agents/$threadId'
path: '/$threadId'
fullPath: '/agents/$threadId'
preLoaderRoute: typeof AgentsThreadIdRouteImport
parentRoute: typeof AgentsRoute
}
'/admin_/evals': {
id: '/admin_/evals'
path: '/admin/evals'
fullPath: '/admin/evals'
preLoaderRoute: typeof AdminEvalsRouteImport
parentRoute: typeof rootRouteImport
}
'/agents/reviews/': {
id: '/agents/reviews/'
path: '/reviews'
fullPath: '/agents/reviews/'
preLoaderRoute: typeof AgentsReviewsIndexRouteImport
parentRoute: typeof AgentsRoute
}
'/agents/automations/': {
id: '/agents/automations/'
path: '/automations'
fullPath: '/agents/automations/'
preLoaderRoute: typeof AgentsAutomationsIndexRouteImport
parentRoute: typeof AgentsRoute
}
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
'/review_/repositories/$owner': {
id: '/review_/repositories/$owner'
path: '/review/repositories/$owner'
fullPath: '/review/repositories/$owner'
preLoaderRoute: typeof ReviewRepositoriesOwnerRouteImport
parentRoute: typeof rootRouteImport
}
'/agents/automations/new': {
id: '/agents/automations/new'
path: '/automations/new'
fullPath: '/agents/automations/new'
preLoaderRoute: typeof AgentsAutomationsNewRouteImport
parentRoute: typeof AgentsRoute
}
'/agents/automations/$scheduleId': {
id: '/agents/automations/$scheduleId'
path: '/automations/$scheduleId'
fullPath: '/agents/automations/$scheduleId'
preLoaderRoute: typeof AgentsAutomationsScheduleIdRouteImport
parentRoute: typeof AgentsRoute
}
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
'/agents/$threadId_/plan': {
id: '/agents/$threadId_/plan'
path: '/$threadId/plan'
fullPath: '/agents/$threadId/plan'
preLoaderRoute: typeof AgentsThreadIdPlanRouteImport
parentRoute: typeof AgentsRoute
}
'/agents/reviews/$owner/$repo/$number': {
id: '/agents/reviews/$owner/$repo/$number'
path: '/reviews/$owner/$repo/$number'
fullPath: '/agents/reviews/$owner/$repo/$number'
preLoaderRoute: typeof AgentsReviewsOwnerRepoNumberRouteImport
parentRoute: typeof AgentsRoute
}
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
}
}
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
interface AgentsRouteChildren {
AgentsThreadIdRoute: typeof AgentsThreadIdRoute
AgentsThreadsRoute: typeof AgentsThreadsRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
AgentsIndexRoute: typeof AgentsIndexRoute
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
AgentsThreadIdPlanRoute: typeof AgentsThreadIdPlanRoute
AgentsAutomationsScheduleIdRoute: typeof AgentsAutomationsScheduleIdRoute
AgentsAutomationsNewRoute: typeof AgentsAutomationsNewRoute
AgentsAutomationsIndexRoute: typeof AgentsAutomationsIndexRoute
AgentsReviewsIndexRoute: typeof AgentsReviewsIndexRoute
AgentsReviewsOwnerRepoNumberRoute: typeof AgentsReviewsOwnerRepoNumberRoute
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
}
const AgentsRouteChildren: AgentsRouteChildren = {
AgentsThreadIdRoute: AgentsThreadIdRoute,
AgentsThreadsRoute: AgentsThreadsRoute,
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
AgentsIndexRoute: AgentsIndexRoute,
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
AgentsThreadIdPlanRoute: AgentsThreadIdPlanRoute,
AgentsAutomationsScheduleIdRoute: AgentsAutomationsScheduleIdRoute,
AgentsAutomationsNewRoute: AgentsAutomationsNewRoute,
AgentsAutomationsIndexRoute: AgentsAutomationsIndexRoute,
AgentsReviewsIndexRoute: AgentsReviewsIndexRoute,
AgentsReviewsOwnerRepoNumberRoute: AgentsReviewsOwnerRepoNumberRoute,
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
}
const AgentsRouteWithChildren =
AgentsRoute._addFileChildren(AgentsRouteChildren)
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
const rootRouteChildren: RootRouteChildren = {
IndexRoute: IndexRoute,
AdminRoute: AdminRoute,
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
AgentsRoute: AgentsRouteWithChildren,
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
CloudAgentsRoute: CloudAgentsRoute,
IntegrationsRoute: IntegrationsRoute,
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
LoginRoute: LoginRoute,
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
MySettingsRoute: MySettingsRoute,
ReviewRoute: ReviewRoute,
UsageRoute: UsageRoute,
AdminEvalsRoute: AdminEvalsRoute,
AgentsInstructionsRoute: AgentsInstructionsRoute,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
AgentsSnapshotsRoute: AgentsSnapshotsRoute,
ReviewStylesRoute: ReviewStylesRoute,
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
ReviewRepositoriesOwnerRoute: ReviewRepositoriesOwnerRoute,
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
}
export const routeTree = rootRouteImport
._addFileChildren(rootRouteChildren)
._addFileTypes<FileRouteTypes>()
import type { getRouter } from './router.tsx'
import type { createStart } from '@tanstack/react-start'
declare module '@tanstack/react-start' {
interface Register {
ssr: true
router: Awaited<ReturnType<typeof getRouter>>
}
}