open-swe/agent/dashboard/team_settings.py

528 lines
22 KiB
Python
Raw Permalink Normal View History

feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
"""Team-wide Open SWE Review (Bugbot) settings stored in LangGraph Store.
A single record keyed ``"default"`` keeps all instance-wide reviewer
configuration in one place. Per-repo style prompts live in
:mod:`agent.dashboard.review_styles`.
"""
from __future__ import annotations
import logging
import os
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
from datetime import UTC, datetime
from typing import Any, Literal
from langgraph_sdk import get_client
from pydantic import BaseModel, field_validator, model_validator
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155) * feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) Ports four upstream commits that add opt-in LLM call routing through the LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs, no-agent-attribution, bun toolchain). - #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings toggle, admin UI section, wired into make_model for all graph entrypoints - #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over platform LANGSMITH_API_KEY - #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host, SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware - #1678 (73b7d1c0): fix OpenAI Responses reasoning replay — SanitizeOpenAIResponsesMiddleware, store/include config for encrypted reasoning content, reasoning_effort coercion for Chat Completions fallback Refs #134 * fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity - Downgrade logger.warning to logger.debug in gateway_overrides for not-routed providers and missing API key (Bedrock is the default provider in this fork, so these are expected steady states) - Add Bedrock to the LLMGatewaySection route-toggle description so admins know it is not routed through the gateway - Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with server.py and reviewer.py - Restore the Bedrock region comment in model.py that explains the AWS_REGION / AWS_DEFAULT_REGION precedence Refs #138 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-09 14:44:15 -04:00
from ..utils.gateway import resolve_gateway_enabled
from .options import (
FABLE_MODEL_IDS,
SUPPORTED_MODEL_IDS,
default_model_pair,
gate_fable_model,
model_supports_effort,
provider_fallback_pair,
)
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
logger = logging.getLogger(__name__)
TEAM_SETTINGS_NAMESPACE: list[str] = ["team_settings"]
TEAM_SETTINGS_KEY = "default"
# Cap the org-wide guidelines so a runaway value can't dominate the reviewer
# prompt. Generous enough for a detailed policy, small enough to stay bounded.
ORG_GUIDELINES_MAX_CHARS = 10_000
chore: sync upstream/main, defer #1621 modular webhooks (#81) * chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
REVIEW_TRACING_PROJECT_MAX_CHARS = 256
feat: distill Sea Haven conventions into agent prompt, reviewer, and fork docs (#113) * Add Dependabot ignore for @types/node semver-major bumps Prevent Dependabot from proposing wrong-direction @types/node major bumps (e.g. 24 -> 26). /ui runs on Node 24 on Vercel; a too-new types major still compiles but describes APIs absent at runtime. Refs: #110 * feat(agent): seed all-repos custom instructions in default_prompt.md Distill the universally-applicable Sea Haven authoring conventions into the team-default Custom Instructions the main agent gets on every repo: secrets/ config placement, keep-docs-in-sync, verify-before-push, re-run-real-gates after delegating, confirm-a-convention-before-adopting, and house writing style. Toolchain references are generalized (not tied to a specific stack). * feat(reviewer): seed Sea Haven review baseline as org-guidelines default Bake DEFAULT_ORG_REVIEW_GUIDELINES (severity model, secrets, security surface, tests, naming, deferred-work-needs-an-issue) and default org_guidelines to it in _default_settings(). The reviewer now applies the Sea Haven baseline on every repo until a workspace admin overrides it with a non-empty value via the dashboard. Stack-agnostic and well under the 10k-char cap. * refactor(prompt): consolidate duplicated COMMIT_PR_SECTION + add fork-sync runbook COMMIT_PR_SECTION had two overlapping passes with a contradictory PR-title rule (a fixed 'type: description' form vs the repo-aware detection). Collapse into one numbered sequence (lint -> commit -> push/PR -> notify), keep the authoritative repo-aware title rule, and drop the duplicate notify step. All IMPORTANT directives (force-push ban, workflow-approval, autonomy, 403 handling) are preserved verbatim. Add a fork-maintenance runbook to CLAUDE.md distilling the durable upstream-sync methodology (conflict triage, deferred-refactor resolution rule, the silent re-import/wiring hazards, test-impl-same-side, layered CI). * refactor(prompt): adopt conventional-commit style Flip the Sea Haven authoring convention baked into the agent prompt from imperative/no-prefix to conventional-commit style: - Commit subjects and the no-gate PR-title default now use type(scope): description with the allowed type set (feat, fix, docs, style, refactor, perf, test, build, ci, chore, revert, release). - Branch prefixes expanded to feature/, fix/, hotfix/, chore/, docs/, refactor/, release/ (kebab-case description). - The repo-aware gate detection is preserved: a repo's own title gate still wins and may narrow the allowed types/scopes. Updated test_github_comment_prompts.py to assert the new convention.
2026-07-02 18:29:15 -04:00
# Sea Haven review baseline seeded as the org-wide guidelines default. Surfaces
# in the reviewer prompt for every repo until an admin overrides it with a
# non-empty value via the dashboard (PUT /team-settings). Keep it stack-agnostic
# and well under ORG_GUIDELINES_MAX_CHARS.
DEFAULT_ORG_REVIEW_GUIDELINES = """\
Sea Haven review baseline (applies to every repo unless a repo-specific guideline overrides it):
- Severity: map findings to critical / high / medium / low. Reserve critical and high for correctness bugs, security issues, data loss, or broken contracts — not style.
- Secrets & config: flag any hardcoded secret, credential, or real `.env`/config value committed to source, and any sensitive value placed outside the platform's secrets manager.
- Security surface (raise as high): changes to authentication/authorization or access checks; IAM/policy/permission or infrastructure-access changes; changes to the exported signature or contract of a public handler/endpoint; and untrusted-input handling (request parsing, deserialization, file uploads, SSRF-prone fetches, and template/SQL/command construction — flag unescaped interpolation of dynamic or user-controlled data).
- Tests: flag new behavior that ships without a corresponding test, and "fixes" that only silence a check (added excludes, `noqa` / `# type: ignore`, skipped or `xfail`ed tests).
- Naming & conventions: flag resources or code that break the repo's established naming and layout conventions.
- Deferred work: a finding the author chooses to defer must be captured in a tracked issue, not dropped silently.
Only file a finding that anchors to a changed line and names a concrete failure mode. Do not police pre-existing issues outside the diff or raise pure style nits."""
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
class TeamSettingsUpdate(BaseModel):
auto_verdict: bool = False
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
review_draft_prs: bool = False
pr_summaries: bool = True
review_trace_links: bool = True
feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155) * feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) Ports four upstream commits that add opt-in LLM call routing through the LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs, no-agent-attribution, bun toolchain). - #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings toggle, admin UI section, wired into make_model for all graph entrypoints - #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over platform LANGSMITH_API_KEY - #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host, SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware - #1678 (73b7d1c0): fix OpenAI Responses reasoning replay — SanitizeOpenAIResponsesMiddleware, store/include config for encrypted reasoning content, reasoning_effort coercion for Chat Completions fallback Refs #134 * fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity - Downgrade logger.warning to logger.debug in gateway_overrides for not-routed providers and missing API key (Bedrock is the default provider in this fork, so these are expected steady states) - Add Bedrock to the LLMGatewaySection route-toggle description so admins know it is not routed through the gateway - Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with server.py and reviewer.py - Restore the Bedrock region comment in model.py that explains the AWS_REGION / AWS_DEFAULT_REGION precedence Refs #138 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-09 14:44:15 -04:00
# Tri-state LLM Gateway toggle: True/False is authoritative, None inherits the
# LANGSMITH_GATEWAY_ENABLED deployment default.
gateway_enabled: bool | None = None
fable_enabled: bool = False
chore: sync upstream/main, defer #1621 modular webhooks (#81) * chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
review_tracing_project: str | None = None
org_guidelines: str | None = None
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
default_agent_model: str | None = None
default_agent_reasoning_effort: str | None = None
default_agent_subagent_model: str | None = None
default_agent_subagent_reasoning_effort: str | None = None
default_repo: str | None = None
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
default_reviewer_model: str | None = None
default_reviewer_reasoning_effort: str | None = None
default_reviewer_subagent_model: str | None = None
default_reviewer_subagent_reasoning_effort: str | None = None
feat: AI-sorted PR review view with diff grouping (#1544) * feat: AI-sorted PR review view with diff grouping Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale. Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it. * feat(reviews): richer AI-sorted explanations + sidebar polish Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link. Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path. * fix(reviews): drop stale diff groups from the AI-sorted view When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 13:59:37 -07:00
default_grouping_model: str | None = None
default_grouping_reasoning_effort: str | None = None
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
default_chat_model: str | None = None
default_chat_reasoning_effort: str | None = None
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
@field_validator("org_guidelines", mode="before")
@classmethod
def _normalize_org_guidelines(cls, v: object) -> str | None:
if v is None:
return None
if not isinstance(v, str):
raise ValueError("org_guidelines must be a string")
text = v.strip()
if not text:
return None
if len(text) > ORG_GUIDELINES_MAX_CHARS:
raise ValueError(
f"org_guidelines must be at most {ORG_GUIDELINES_MAX_CHARS} characters"
)
return text
chore: sync upstream/main, defer #1621 modular webhooks (#81) * chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
@field_validator("review_tracing_project", mode="before")
@classmethod
def _normalize_review_tracing_project(cls, v: object) -> str | None:
if v is None:
return None
if not isinstance(v, str):
raise ValueError("review_tracing_project must be a string")
text = v.strip()
if not text:
return None
if len(text) > REVIEW_TRACING_PROJECT_MAX_CHARS:
raise ValueError(
"review_tracing_project must be at most "
f"{REVIEW_TRACING_PROJECT_MAX_CHARS} characters"
)
return text
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
@model_validator(mode="after")
def _validate_model_pairs(self) -> TeamSettingsUpdate:
self.default_agent_model, self.default_agent_reasoning_effort = _normalize_stale_model_pair(
self.default_agent_model,
self.default_agent_reasoning_effort,
)
self.default_agent_subagent_model, self.default_agent_subagent_reasoning_effort = (
_normalize_stale_model_pair(
self.default_agent_subagent_model,
self.default_agent_subagent_reasoning_effort,
)
)
self.default_reviewer_model, self.default_reviewer_reasoning_effort = (
_normalize_stale_model_pair(
self.default_reviewer_model,
self.default_reviewer_reasoning_effort,
)
)
(
self.default_reviewer_subagent_model,
self.default_reviewer_subagent_reasoning_effort,
) = _normalize_stale_model_pair(
self.default_reviewer_subagent_model,
self.default_reviewer_subagent_reasoning_effort,
)
self.default_grouping_model, self.default_grouping_reasoning_effort = (
_normalize_stale_model_pair(
self.default_grouping_model,
self.default_grouping_reasoning_effort,
)
)
self.default_chat_model, self.default_chat_reasoning_effort = _normalize_stale_model_pair(
self.default_chat_model,
self.default_chat_reasoning_effort,
)
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
_validate_model_effort_pair(
self.default_agent_model, self.default_agent_reasoning_effort, "agent"
)
_validate_model_effort_pair(
self.default_agent_subagent_model,
self.default_agent_subagent_reasoning_effort,
"agent subagent",
)
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
_validate_model_effort_pair(
self.default_reviewer_model, self.default_reviewer_reasoning_effort, "reviewer"
)
_validate_model_effort_pair(
self.default_reviewer_subagent_model,
self.default_reviewer_subagent_reasoning_effort,
"reviewer subagent",
)
feat: AI-sorted PR review view with diff grouping (#1544) * feat: AI-sorted PR review view with diff grouping Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale. Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it. * feat(reviews): richer AI-sorted explanations + sidebar polish Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link. Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path. * fix(reviews): drop stale diff groups from the AI-sorted view When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 13:59:37 -07:00
_validate_model_effort_pair(
self.default_grouping_model,
self.default_grouping_reasoning_effort,
"review diff grouping",
)
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
_validate_model_effort_pair(
self.default_chat_model, self.default_chat_reasoning_effort, "review chat"
)
if not self.fable_enabled:
# Disabling Fable is the provider-data-share kill switch and must always succeed: rather
# than reject a payload that still carries a Fable default, swap each
# Fable default to its safe non-Fable fallback (mirrors the runtime
# gate_fable_model guard) so the stored record can't advertise Fable.
for model_field, effort_field in (
("default_agent_model", "default_agent_reasoning_effort"),
("default_agent_subagent_model", "default_agent_subagent_reasoning_effort"),
("default_reviewer_model", "default_reviewer_reasoning_effort"),
("default_reviewer_subagent_model", "default_reviewer_subagent_reasoning_effort"),
("default_grouping_model", "default_grouping_reasoning_effort"),
("default_chat_model", "default_chat_reasoning_effort"),
):
model = getattr(self, model_field)
if model in FABLE_MODEL_IDS:
new_model, new_effort = gate_fable_model(
model, getattr(self, effort_field), fable_enabled=False
)
setattr(self, model_field, new_model)
setattr(self, effort_field, new_effort)
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
return self
def _validate_model_effort_pair(model: str | None, effort: str | None, role: str) -> None:
if model is None and effort is None:
return
if model is None:
raise ValueError(f"{role} reasoning effort set without a model")
if model not in SUPPORTED_MODEL_IDS:
raise ValueError(f"unsupported {role} model: {model}")
if effort is None or not model_supports_effort(model, effort):
raise ValueError(f"effort {effort!r} not supported by {role} model {model!r}")
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
# Model ids retired from SUPPORTED_MODELS mapped to their current successors,
# so stored admin defaults stay saveable/readable across model upgrades. Direct
# anthropic: ids moved to the Bedrock inference profile in the provider
# migration (#62); OpenAI/Google were dropped entirely and take the global
# default. anthropic:claude-fable-5 deliberately maps to Opus, not Bedrock
# Fable: Fable is gated behind the provider-data-share opt-in and a stale
# record must never resurface it.
_RETIRED_MODEL_REPLACEMENTS: dict[str, str] = {
"anthropic:claude-opus-4-7": "bedrock_converse:us.anthropic.claude-opus-4-8",
"anthropic:claude-opus-4-8": "bedrock_converse:us.anthropic.claude-opus-4-8",
"anthropic:claude-fable-5": "bedrock_converse:us.anthropic.claude-opus-4-8",
"openai:gpt-5.5": "bedrock_converse:us.anthropic.claude-opus-4-8",
"google_genai:gemini-3.5-flash": "bedrock_converse:us.anthropic.claude-opus-4-8",
}
def _normalize_stale_model_pair(
model: str | None, effort: str | None
) -> tuple[str | None, str | None]:
if model is None:
return model, effort
return _RETIRED_MODEL_REPLACEMENTS.get(model, model), effort
_MODEL_PAIR_FIELDS: tuple[tuple[str, str], ...] = (
("default_agent_model", "default_agent_reasoning_effort"),
("default_agent_subagent_model", "default_agent_subagent_reasoning_effort"),
("default_reviewer_model", "default_reviewer_reasoning_effort"),
("default_reviewer_subagent_model", "default_reviewer_subagent_reasoning_effort"),
("default_grouping_model", "default_grouping_reasoning_effort"),
("default_chat_model", "default_chat_reasoning_effort"),
)
def normalize_team_settings_for_response(settings: dict[str, Any]) -> dict[str, Any]:
value = dict(settings)
for model_field, effort_field in _MODEL_PAIR_FIELDS:
model = value.get(model_field)
effort = value.get(effort_field)
if isinstance(model, str):
value[model_field], value[effort_field] = _normalize_stale_model_pair(
model,
effort if isinstance(effort, str) else None,
)
return value
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
def _client():
return get_client()
def _env_default_repo() -> str | None:
owner = os.environ.get("DEFAULT_REPO_OWNER", "").strip()
name = os.environ.get("DEFAULT_REPO_NAME", "").strip()
return f"{owner}/{name}" if owner and name else None
def _parse_repo(value: object) -> dict[str, str] | None:
if not isinstance(value, str):
return None
owner, sep, name = value.strip().partition("/")
if not sep or not owner.strip() or not name.strip():
return None
return {"owner": owner.strip(), "name": name.strip()}
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
def _default_settings() -> dict[str, Any]:
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
fallback_model, fallback_effort = default_model_pair()
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
return {
"auto_verdict": False,
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
"review_draft_prs": False,
"pr_summaries": True,
"review_trace_links": True,
feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155) * feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) Ports four upstream commits that add opt-in LLM call routing through the LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs, no-agent-attribution, bun toolchain). - #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings toggle, admin UI section, wired into make_model for all graph entrypoints - #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over platform LANGSMITH_API_KEY - #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host, SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware - #1678 (73b7d1c0): fix OpenAI Responses reasoning replay — SanitizeOpenAIResponsesMiddleware, store/include config for encrypted reasoning content, reasoning_effort coercion for Chat Completions fallback Refs #134 * fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity - Downgrade logger.warning to logger.debug in gateway_overrides for not-routed providers and missing API key (Bedrock is the default provider in this fork, so these are expected steady states) - Add Bedrock to the LLMGatewaySection route-toggle description so admins know it is not routed through the gateway - Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with server.py and reviewer.py - Restore the Bedrock region comment in model.py that explains the AWS_REGION / AWS_DEFAULT_REGION precedence Refs #138 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-09 14:44:15 -04:00
"gateway_enabled": None,
"fable_enabled": False,
chore: sync upstream/main, defer #1621 modular webhooks (#81) * chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
"review_tracing_project": None,
feat: distill Sea Haven conventions into agent prompt, reviewer, and fork docs (#113) * Add Dependabot ignore for @types/node semver-major bumps Prevent Dependabot from proposing wrong-direction @types/node major bumps (e.g. 24 -> 26). /ui runs on Node 24 on Vercel; a too-new types major still compiles but describes APIs absent at runtime. Refs: #110 * feat(agent): seed all-repos custom instructions in default_prompt.md Distill the universally-applicable Sea Haven authoring conventions into the team-default Custom Instructions the main agent gets on every repo: secrets/ config placement, keep-docs-in-sync, verify-before-push, re-run-real-gates after delegating, confirm-a-convention-before-adopting, and house writing style. Toolchain references are generalized (not tied to a specific stack). * feat(reviewer): seed Sea Haven review baseline as org-guidelines default Bake DEFAULT_ORG_REVIEW_GUIDELINES (severity model, secrets, security surface, tests, naming, deferred-work-needs-an-issue) and default org_guidelines to it in _default_settings(). The reviewer now applies the Sea Haven baseline on every repo until a workspace admin overrides it with a non-empty value via the dashboard. Stack-agnostic and well under the 10k-char cap. * refactor(prompt): consolidate duplicated COMMIT_PR_SECTION + add fork-sync runbook COMMIT_PR_SECTION had two overlapping passes with a contradictory PR-title rule (a fixed 'type: description' form vs the repo-aware detection). Collapse into one numbered sequence (lint -> commit -> push/PR -> notify), keep the authoritative repo-aware title rule, and drop the duplicate notify step. All IMPORTANT directives (force-push ban, workflow-approval, autonomy, 403 handling) are preserved verbatim. Add a fork-maintenance runbook to CLAUDE.md distilling the durable upstream-sync methodology (conflict triage, deferred-refactor resolution rule, the silent re-import/wiring hazards, test-impl-same-side, layered CI). * refactor(prompt): adopt conventional-commit style Flip the Sea Haven authoring convention baked into the agent prompt from imperative/no-prefix to conventional-commit style: - Commit subjects and the no-gate PR-title default now use type(scope): description with the allowed type set (feat, fix, docs, style, refactor, perf, test, build, ci, chore, revert, release). - Branch prefixes expanded to feature/, fix/, hotfix/, chore/, docs/, refactor/, release/ (kebab-case description). - The repo-aware gate detection is preserved: a repo's own title gate still wins and may narrow the allowed types/scopes. Updated test_github_comment_prompts.py to assert the new convention.
2026-07-02 18:29:15 -04:00
"org_guidelines": DEFAULT_ORG_REVIEW_GUIDELINES,
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
"default_agent_model": fallback_model,
"default_agent_reasoning_effort": fallback_effort,
"default_agent_subagent_model": fallback_model,
"default_agent_subagent_reasoning_effort": fallback_effort,
"default_repo": _env_default_repo(),
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
"default_reviewer_model": fallback_model,
"default_reviewer_reasoning_effort": fallback_effort,
"default_reviewer_subagent_model": fallback_model,
"default_reviewer_subagent_reasoning_effort": fallback_effort,
feat: AI-sorted PR review view with diff grouping (#1544) * feat: AI-sorted PR review view with diff grouping Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale. Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it. * feat(reviews): richer AI-sorted explanations + sidebar polish Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link. Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path. * fix(reviews): drop stale diff groups from the AI-sorted view When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 13:59:37 -07:00
# No hardcoded grouping default: unset means "inherit the Reviewer
# subagent default".
"default_grouping_model": None,
"default_grouping_reasoning_effort": None,
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
# No hardcoded chat default: unset means "inherit the Agent default".
"default_chat_model": None,
"default_chat_reasoning_effort": None,
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
"updated_at": None,
}
async def get_team_settings() -> dict[str, Any]:
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
defaults = _default_settings()
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
try:
item = await _client().store.get_item(TEAM_SETTINGS_NAMESPACE, TEAM_SETTINGS_KEY)
except Exception as e:
logger.debug("team settings lookup failed: %s", e)
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
return defaults
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
if item is None:
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
return defaults
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
value = item.get("value") if isinstance(item, dict) else getattr(item, "value", None)
if not isinstance(value, dict):
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
return defaults
# Skip None-valued model fields so legacy records (or PUTs that cleared the
# selection) still surface the hardcoded default instead of a null.
overlay = {k: v for k, v in value.items() if v is not None}
merged = {**defaults, **overlay}
feat: activate PR babysitting UI toggles for autofix and trigger mode (#1561) * feat: activate PR babysitting UI toggles for autofix and trigger mode Remove the "coming soon" gating on the Autofix Mode, Autofix Severity Threshold, and Trigger Mode controls in the review settings page so admins can enable CI auto-fix and review-comment resolution on PRs that Open SWE opens. The backend (ci_autofix.py, webapp.py webhook routing) was already fully wired — only the UI was disabled. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: simplify autofix to on/off toggle, remove severity threshold Replace the four-level AutofixMode (off/low/medium/high) and the autofix_severity_threshold setting with a single boolean autofix_enabled toggle. The severity threshold was leftover from the reviewer finding-severity model and does not apply to CI autofix; the agent should fix any failing CI and resolve any comments on PRs it opens. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: move autofix toggle to per-user profile, remove team-level setting The autofix toggle is now per-user (auto_fix_ci in the user profile) instead of team-level (admin-only). This uses the existing auto_fix_ci field that was already in ProfileUpdate but never wired up. Changes: - ci_autofix.py: check per-user auto_fix_ci profile flag after resolving the agent thread's github_login, instead of checking team-level autofix_enabled before knowing the PR - webapp.py: removed early is_autofix_enabled() webhook gates; the per-user check now happens in ci_autofix.py once the thread is found - team_settings.py: removed autofix_enabled field, is_autofix_enabled() - cloud-agents.tsx: enabled the auto_fix_ci toggle (was comingSoon) - review.tsx: removed the admin-level autofix switch - Updated tests and AGENTS.md The agent graph (not the reviewer) is what gets dispatched - this was already correct in ci_autofix.py line 223: client.runs.create( thread_id, "agent", ...). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: batch PR babysitting events Remove the leftover trigger-mode gate from PR babysitting and batch new CI/review events while an agent run is already active so the running agent can handle the latest PR state before finishing. Also moves review-feedback permission checks behind the per-user opt-out and applies the auto-fix profile gate to merge-conflict babysitting. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: consume batched babysitting events Teach the agent queue middleware to turn pending PR babysitting metadata into an injected instruction for the active run, so batched CI/review events are not dropped while still avoiding duplicate run creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review findings in PR babysitting batching - Route batched events through the LangGraph store (read in-process by the message-queue middleware) instead of a per-model-call threads.get on every agent thread. - Only record an attempt / mark the head SHA handled on a real dispatch, not on a batch, so an event isn't permanently dropped if the in-flight run ends before consuming it. - Carry the reviewer's comment through batched review feedback instead of replacing it with a generic re-check nudge. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 14:12:04 -07:00
for stale_field in (
"trigger_mode",
"autofix_mode",
"autofix_severity_threshold",
"autofix_enabled",
chore: sync upstream/main, defer #1621 modular webhooks (#81) * chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
"review_author_context_enabled",
feat: activate PR babysitting UI toggles for autofix and trigger mode (#1561) * feat: activate PR babysitting UI toggles for autofix and trigger mode Remove the "coming soon" gating on the Autofix Mode, Autofix Severity Threshold, and Trigger Mode controls in the review settings page so admins can enable CI auto-fix and review-comment resolution on PRs that Open SWE opens. The backend (ci_autofix.py, webapp.py webhook routing) was already fully wired — only the UI was disabled. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: simplify autofix to on/off toggle, remove severity threshold Replace the four-level AutofixMode (off/low/medium/high) and the autofix_severity_threshold setting with a single boolean autofix_enabled toggle. The severity threshold was leftover from the reviewer finding-severity model and does not apply to CI autofix; the agent should fix any failing CI and resolve any comments on PRs it opens. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: move autofix toggle to per-user profile, remove team-level setting The autofix toggle is now per-user (auto_fix_ci in the user profile) instead of team-level (admin-only). This uses the existing auto_fix_ci field that was already in ProfileUpdate but never wired up. Changes: - ci_autofix.py: check per-user auto_fix_ci profile flag after resolving the agent thread's github_login, instead of checking team-level autofix_enabled before knowing the PR - webapp.py: removed early is_autofix_enabled() webhook gates; the per-user check now happens in ci_autofix.py once the thread is found - team_settings.py: removed autofix_enabled field, is_autofix_enabled() - cloud-agents.tsx: enabled the auto_fix_ci toggle (was comingSoon) - review.tsx: removed the admin-level autofix switch - Updated tests and AGENTS.md The agent graph (not the reviewer) is what gets dispatched - this was already correct in ci_autofix.py line 223: client.runs.create( thread_id, "agent", ...). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: batch PR babysitting events Remove the leftover trigger-mode gate from PR babysitting and batch new CI/review events while an agent run is already active so the running agent can handle the latest PR state before finishing. Also moves review-feedback permission checks behind the per-user opt-out and applies the auto-fix profile gate to merge-conflict babysitting. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: consume batched babysitting events Teach the agent queue middleware to turn pending PR babysitting metadata into an injected instruction for the active run, so batched CI/review events are not dropped while still avoiding duplicate run creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review findings in PR babysitting batching - Route batched events through the LangGraph store (read in-process by the message-queue middleware) instead of a per-model-call threads.get on every agent thread. - Only record an attempt / mark the head SHA handled on a real dispatch, not on a batch, so an event isn't permanently dropped if the in-flight run ends before consuming it. - Carry the reviewer's comment through batched review feedback instead of replacing it with a generic re-check nudge. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 14:12:04 -07:00
):
merged.pop(stale_field, None)
return normalize_team_settings_for_response(merged)
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
async def upsert_team_settings(update: TeamSettingsUpdate) -> dict[str, Any]:
value: dict[str, Any] = {
"auto_verdict": update.auto_verdict,
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
"review_draft_prs": update.review_draft_prs,
"pr_summaries": update.pr_summaries,
"review_trace_links": update.review_trace_links,
feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155) * feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) Ports four upstream commits that add opt-in LLM call routing through the LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs, no-agent-attribution, bun toolchain). - #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings toggle, admin UI section, wired into make_model for all graph entrypoints - #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over platform LANGSMITH_API_KEY - #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host, SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware - #1678 (73b7d1c0): fix OpenAI Responses reasoning replay — SanitizeOpenAIResponsesMiddleware, store/include config for encrypted reasoning content, reasoning_effort coercion for Chat Completions fallback Refs #134 * fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity - Downgrade logger.warning to logger.debug in gateway_overrides for not-routed providers and missing API key (Bedrock is the default provider in this fork, so these are expected steady states) - Add Bedrock to the LLMGatewaySection route-toggle description so admins know it is not routed through the gateway - Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with server.py and reviewer.py - Restore the Bedrock region comment in model.py that explains the AWS_REGION / AWS_DEFAULT_REGION precedence Refs #138 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-09 14:44:15 -04:00
"gateway_enabled": update.gateway_enabled,
"fable_enabled": update.fable_enabled,
chore: sync upstream/main, defer #1621 modular webhooks (#81) * chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
"review_tracing_project": update.review_tracing_project,
"org_guidelines": update.org_guidelines,
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
"default_agent_model": update.default_agent_model,
"default_agent_reasoning_effort": update.default_agent_reasoning_effort,
"default_agent_subagent_model": update.default_agent_subagent_model,
"default_agent_subagent_reasoning_effort": update.default_agent_subagent_reasoning_effort,
"default_repo": update.default_repo,
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
"default_reviewer_model": update.default_reviewer_model,
"default_reviewer_reasoning_effort": update.default_reviewer_reasoning_effort,
"default_reviewer_subagent_model": update.default_reviewer_subagent_model,
"default_reviewer_subagent_reasoning_effort": update.default_reviewer_subagent_reasoning_effort,
feat: AI-sorted PR review view with diff grouping (#1544) * feat: AI-sorted PR review view with diff grouping Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale. Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it. * feat(reviews): richer AI-sorted explanations + sidebar polish Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link. Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path. * fix(reviews): drop stale diff groups from the AI-sorted view When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 13:59:37 -07:00
"default_grouping_model": update.default_grouping_model,
"default_grouping_reasoning_effort": update.default_grouping_reasoning_effort,
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
"default_chat_model": update.default_chat_model,
"default_chat_reasoning_effort": update.default_chat_reasoning_effort,
feat: refactor UI into sidebar layout (#1318) * feat(dashboard): refactor UI into Cursor-style sidebar layout Replace the top-bar AppHeader with a left sidebar that holds the four primary sections (My Settings, Cloud Agents, Open SWE Review, Integrations) and the user account menu at the bottom. Admin remains a hidden route reachable from the sidebar only when is_admin. UI: - AppShell + AppSidebar with bottom-anchored avatar/sign-out menu. - My Settings: profile (first/last name) + PR destination preference. - Cloud Agents: model/effort, default repo, base branch, branch prefix, PR toggles (auto-fix CI, create PRs, allow artifacts), Slack notifications. - Open SWE Review: trigger mode, draft reviews, PR summaries, autofix mode + severity threshold, plus the existing per-repo review-style prompts (was /review-styles). - Integrations: GitHub install status + Slack/Linear status rows. - New Switch and Menu primitives. Backend: - Extend ProfileUpdate with first_name, last_name, base_branch, branch_prefix, auto_fix_ci, create_prs, allow_artifacts, slack_notifications, preferred_pr_destination. - Add team_settings module + /team-settings GET (any session) / PUT (admin only) for the reviewer configuration. * fix(dashboard): replace base-ui Menu with plain dropdown, stabilise Selects - base-ui's Menu hit an "Invalid hook call / Cannot read properties of null (reading 'useRef')" crash inside fastComponent under Vite's dep optimisation. The sidebar account menu is the only consumer, so swap it for a useState-driven dropdown with click-outside + Esc handling and drop the unused Menu wrapper. - Initialise the Open SWE Review settings form with the same defaults the server returns instead of `null`, so the Selects don't switch from uncontrolled (`undefined`) to controlled on first data load. * fix(dashboard): address review feedback on admin overwrite + form reset - admin PUT /admin/profiles/{login} only sent model/effort/repo from the admin form. ProfileUpdate's Pydantic defaults then wrote auto_fix_ci / create_prs / slack_notifications etc back onto the target user's profile, clobbering whatever they had configured. Switch to model_dump(exclude_unset=True) and overlay the incoming fields on the user's existing stored profile so only sent fields change. - my-settings.tsx and cloud-agents.tsx initialised local state from profile.data on every change. Because useSaveProfile updates the cached profile on success, toggling any control would overwrite the user's typed-but-unsaved input in another field. Guard the initialisation effect with a ref so it runs once on first load. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-20 14:13:23 -07:00
"updated_at": datetime.now(UTC).isoformat(),
}
await _client().store.put_item(TEAM_SETTINGS_NAMESPACE, TEAM_SETTINGS_KEY, value)
return value
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
async def get_team_default_repo() -> dict[str, str] | None:
settings = await get_team_settings()
return _parse_repo(settings.get("default_repo"))
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
async def get_team_default_model(
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
role: Literal["agent", "reviewer", "chat"],
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
) -> tuple[str, str]:
"""Return the team-wide default ``(model_id, reasoning_effort)`` for ``role``.
Always returns a valid pair, resolved in order: the admin-configured pair if
still supported; otherwise the newest supported model for the same provider
(so a stale Anthropic/OpenAI selection stays on its provider rather than
jumping cross-provider); otherwise the hardcoded global default from
:func:`agent.dashboard.options.default_model_pair`.
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
``"chat"`` (the review-page PR chat) has no hardcoded default: when its
admin setting is unset/invalid it inherits the team **agent** default.
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
"""
settings = await get_team_settings()
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
if role == "chat":
model = settings.get("default_chat_model")
effort = settings.get("default_chat_reasoning_effort")
if (
isinstance(model, str)
and isinstance(effort, str)
and model in SUPPORTED_MODEL_IDS
and model_supports_effort(model, effort)
):
return _resolve_default_pair(model, effort)
# Inherit the Agent default when no chat-specific model is configured.
model = settings.get("default_agent_model")
effort = settings.get("default_agent_reasoning_effort")
elif role == "agent":
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
model = settings.get("default_agent_model")
effort = settings.get("default_agent_reasoning_effort")
else:
model = settings.get("default_reviewer_model")
effort = settings.get("default_reviewer_reasoning_effort")
return _resolve_default_pair(model, effort)
async def get_team_default_model_pair(
role: Literal["agent", "reviewer"],
) -> tuple[tuple[str, str], tuple[str, str]]:
"""Return default ``(main, subagent)`` model pairs for ``role`` from one store read."""
settings = await get_team_settings()
if role == "agent":
main = _resolve_default_pair(
settings.get("default_agent_model"),
settings.get("default_agent_reasoning_effort"),
)
subagent = _resolve_default_pair(
settings.get("default_agent_subagent_model"),
settings.get("default_agent_subagent_reasoning_effort"),
)
else:
main = _resolve_default_pair(
settings.get("default_reviewer_model"),
settings.get("default_reviewer_reasoning_effort"),
)
subagent = _resolve_default_pair(
settings.get("default_reviewer_subagent_model"),
settings.get("default_reviewer_subagent_reasoning_effort"),
)
return main, subagent
feat: AI-sorted PR review view with diff grouping (#1544) * feat: AI-sorted PR review view with diff grouping Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale. Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it. * feat(reviews): richer AI-sorted explanations + sidebar polish Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link. Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path. * fix(reviews): drop stale diff groups from the AI-sorted view When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 13:59:37 -07:00
async def get_team_default_grouping_model() -> tuple[str, str]:
"""Return the team-wide default ``(model_id, reasoning_effort)`` for the
review diff-grouping pass.
When no grouping-specific model is configured (or it's no longer
supported), inherit the team **reviewer subagent** default — the grouping
pass is a cheap, fast companion to the reviewer, so it should track that
cheaper tier rather than the primary reviewer model.
"""
settings = await get_team_settings()
model = settings.get("default_grouping_model")
effort = settings.get("default_grouping_reasoning_effort")
if (
isinstance(model, str)
and isinstance(effort, str)
and model in SUPPORTED_MODEL_IDS
and model_supports_effort(model, effort)
):
return _resolve_default_pair(model, effort)
return _resolve_default_pair(
settings.get("default_reviewer_subagent_model"),
settings.get("default_reviewer_subagent_reasoning_effort"),
)
async def get_team_review_trace_links_enabled() -> bool:
"""Return whether GitHub review bodies should include a LangSmith trace link."""
settings = await get_team_settings()
return bool(settings.get("review_trace_links", True))
async def get_team_auto_verdict_enabled() -> bool:
"""Return whether automatic reviewer verdicts are enabled team-wide."""
settings = await get_team_settings()
value = settings.get("auto_verdict")
return bool(value) if isinstance(value, bool) else False
feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155) * feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) Ports four upstream commits that add opt-in LLM call routing through the LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs, no-agent-attribution, bun toolchain). - #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings toggle, admin UI section, wired into make_model for all graph entrypoints - #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over platform LANGSMITH_API_KEY - #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host, SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware - #1678 (73b7d1c0): fix OpenAI Responses reasoning replay — SanitizeOpenAIResponsesMiddleware, store/include config for encrypted reasoning content, reasoning_effort coercion for Chat Completions fallback Refs #134 * fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity - Downgrade logger.warning to logger.debug in gateway_overrides for not-routed providers and missing API key (Bedrock is the default provider in this fork, so these are expected steady states) - Add Bedrock to the LLMGatewaySection route-toggle description so admins know it is not routed through the gateway - Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with server.py and reviewer.py - Restore the Bedrock region comment in model.py that explains the AWS_REGION / AWS_DEFAULT_REGION precedence Refs #138 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-09 14:44:15 -04:00
async def get_team_gateway_enabled() -> bool | None:
"""Return the stored LLM Gateway toggle (``None`` means inherit the env default)."""
settings = await get_team_settings()
value = settings.get("gateway_enabled")
return value if isinstance(value, bool) else None
async def get_team_fable_enabled() -> bool:
"""Return whether Fable models are enabled for the team."""
settings = await get_team_settings()
value = settings.get("fable_enabled")
return bool(value) if isinstance(value, bool) else False
feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155) * feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) Ports four upstream commits that add opt-in LLM call routing through the LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs, no-agent-attribution, bun toolchain). - #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings toggle, admin UI section, wired into make_model for all graph entrypoints - #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over platform LANGSMITH_API_KEY - #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host, SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware - #1678 (73b7d1c0): fix OpenAI Responses reasoning replay — SanitizeOpenAIResponsesMiddleware, store/include config for encrypted reasoning content, reasoning_effort coercion for Chat Completions fallback Refs #134 * fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity - Downgrade logger.warning to logger.debug in gateway_overrides for not-routed providers and missing API key (Bedrock is the default provider in this fork, so these are expected steady states) - Add Bedrock to the LLMGatewaySection route-toggle description so admins know it is not routed through the gateway - Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with server.py and reviewer.py - Restore the Bedrock region comment in model.py that explains the AWS_REGION / AWS_DEFAULT_REGION precedence Refs #138 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-09 14:44:15 -04:00
async def get_effective_gateway_enabled() -> bool:
"""Resolve whether LLM Gateway routing is on: team setting, else env default."""
return resolve_gateway_enabled(await get_team_gateway_enabled())
chore: sync upstream/main, defer #1621 modular webhooks (#81) * chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
async def get_team_review_tracing_project() -> str | None:
"""Return the LangSmith tracing project used for PR trace resolution."""
settings = await get_team_settings()
value = settings.get("review_tracing_project")
if isinstance(value, str) and value.strip():
return value.strip()
return None
async def get_org_review_guidelines() -> str | None:
"""Return the org-wide reviewer guidelines supplement, if configured."""
settings = await get_team_settings()
value = settings.get("org_guidelines")
if isinstance(value, str) and value.strip():
return value.strip()
return None
async def get_team_default_subagent_model(
role: Literal["agent", "reviewer"],
) -> tuple[str, str]:
"""Return the team-wide default subagent ``(model_id, reasoning_effort)`` for ``role``."""
settings = await get_team_settings()
if role == "agent":
model = settings.get("default_agent_subagent_model")
effort = settings.get("default_agent_subagent_reasoning_effort")
else:
model = settings.get("default_reviewer_subagent_model")
effort = settings.get("default_reviewer_subagent_reasoning_effort")
return _resolve_default_pair(model, effort)
def _resolve_default_pair(model: object, effort: object) -> tuple[str, str]:
"""Supported pair if valid, else same-provider fallback, else global default."""
if (
isinstance(model, str)
and isinstance(effort, str)
and model in SUPPORTED_MODEL_IDS
and model_supports_effort(model, effort)
):
return model, effort
provider_pair = provider_fallback_pair(model, effort)
if provider_pair is not None:
return provider_pair
return default_model_pair()