open-swe/agent/server.py

941 lines
37 KiB
Python
Raw Normal View History

"""Main entry point and CLI loop for Open SWE agent."""
# ruff: noqa: E402
# Suppress deprecation warnings from langchain_core (e.g., Pydantic V1 on Python 3.14+)
# ruff: noqa: E402
import logging
import os
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
import time
import warnings
from collections.abc import Sequence
from typing import Any
logger = logging.getLogger(__name__)
from langgraph.graph.state import RunnableConfig
from langgraph.pregel import Pregel
from langgraph_sdk import get_client
warnings.filterwarnings("ignore", module="langchain_core._api.deprecation")
import asyncio
# Suppress Pydantic v1 compatibility warnings from langchain on Python 3.14+
warnings.filterwarnings("ignore", message=".*Pydantic V1.*", category=UserWarning)
from deepagents import create_deep_agent
from deepagents.backends import LangSmithSandbox
2026-02-10 18:18:04 -08:00
from deepagents.backends.protocol import SandboxBackendProtocol
from deepagents.middleware.subagents import GENERAL_PURPOSE_SUBAGENT, SubAgent
from langchain.agents.middleware import ModelCallLimitMiddleware
from langchain_core.language_models import BaseChatModel
from langsmith.sandbox import SandboxClientError
from .dashboard.admin import is_observability_authorized
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
from .dashboard.agent_overrides import (
load_profile,
normalize_profile_overrides,
normalize_profile_subagent_overrides,
feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60) * feat: default Slack/dashboard/schedule PRs + commits to the app identity (#57) Slack/dashboard/schedule runs now author PRs and run git/gh operations as the GitHub App seahaven-openswe[bot] by default (matching GitHub-issue runs), so the self-review 422 is impossible by construction rather than guarded in the prompt. A profile flag author_prs_as_user restores per-user attribution. - open_pull_request._resolve_pr_author_token + auth.resolve_github_token: default to the installation token for these sources; per-user only when opted in. - authorship: commit identity -> seahaven-openswe[bot] (numeric noreply; accepted Vercel-resolution risk, documented inline). - self-trigger safety: INTERNAL_BOT_LOGINS + webapp/reviewer_reconcile/reply markers recognize seahaven-openswe[bot] (bot-authored events are now ours). Supersedes the prompt-only guard in #58. * fix: author commits as the app bot in the default path (SH-IDSPLIT-01) Security review found the commit identity was NOT actually unified to the bot: resolve_triggering_user_identity got a 403 from the installation token and fell back to configurable['github_login'], so commits were still authored as the triggering user (commit=user, push+PR=bot — a three-way split that missed the stated goal). Now gate the triggering-user identity resolution on the same default-bot decision as the token: slack/dashboard/schedule default to the app bot identity unless author_prs_as_user is set. * docs(security): record AUTHZ-SLACK-BOT-DEFAULT-001 as an accepted residual (#59) Single-user deployment; bounded by App-on-pilot + ALLOWED_GITHUB_REPOS lock. Revisit (add a per-user gate) before expanding users or the App installation.
2026-06-29 14:22:33 -04:00
profile_author_prs_as_user,
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
profile_create_prs,
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
resolve_github_login,
)
from .dashboard.agent_usage import record_agent_thread_usage
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
from .dashboard.options import DEFAULT_MODEL_ID, SUPPORTED_MODEL_IDS, model_supports_effort
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
from .dashboard.repo_snapshots import resolve_repo_snapshot_id
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
from .dashboard.team_settings import (
get_team_default_model_pair,
get_team_default_repo,
)
from .dashboard.user_mappings import email_for_login
from .integrations.corridor_mcp import load_corridor_tools
from .integrations.currents_tools import load_currents_tools
from .integrations.datadog_mcp import load_datadog_tools
from .integrations.langsmith import _configure_github_proxy
from .integrations.langsmith_tools import load_langsmith_tools
from .integrations.notion_mcp import load_notion_tools
from .middleware import (
ModelFallbackMiddleware,
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
PlanModeMiddleware,
RepairOrphanedToolCallsMiddleware,
SandboxCircuitBreakerMiddleware,
SanitizeThinkingBlocksMiddleware,
SanitizeToolInputsMiddleware,
feat: add optional Slack Assistants API typing status indicator (#1269) * feat: add optional Slack Assistants API typing status indicator Mirrors OpenClaw's pragmatic approach: instead of rebuilding around assistant_thread_started events, just opt into assistants.threads.setStatus to show 'is thinking…' while the agent is working, and clear it when post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so it can be toggled without touching code. * fix(slack): drop redundant clear, add status heartbeat across model calls - Slack auto-clears the typing indicator on bot post; remove the explicit assistants.threads.setStatus("") call from post_slack_thread_reply. - The indicator expires after ~2 minutes; add a before_model middleware that refreshes it on every model tick so it stays visible across long agent runs. Reuses the existing slack_thread.{channel_id,thread_ts} configurable already plumbed for notify_step_limit. - chat:write is sufficient on the bot token (assistant:write is on the way out per Slack docs); no scope or app-config change required. * feat(slack): contextual status text + rotating loading_messages - set_slack_assistant_status now accepts an optional loading_messages list (capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus payload so Slack rotates through them client-side. - The heartbeat middleware derives a contextual status from the last assistant message's tool calls (e.g. "searching the codebase…" after grep, "running commands…" after execute), falling back to the default "is thinking…" when no tool calls or unknown tool name. - Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the contextual status on each refresh. * fix slack assistant status lifecycle --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 10:21:55 -07:00
SlackAssistantStatusMiddleware,
ToolArtifactMiddleware,
ToolErrorMiddleware,
check_message_queue_before_model,
notify_step_limit_reached,
refresh_github_proxy_before_model,
)
from .prompt import construct_system_prompt
from .tools import (
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
enter_plan_mode,
fetch_url,
http_request,
linear_comment,
linear_create_issue,
linear_delete_issue,
linear_get_issue,
linear_get_issue_comments,
linear_list_teams,
linear_update_issue,
open_pull_request,
request_pr_review,
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
save_plan,
schedule_thread_wakeup,
slack_read_thread_messages,
slack_thread_reply,
web_search,
)
from .utils.auth import resolve_github_token
from .utils.authorship import (
OPEN_SWE_BOT_EMAIL,
OPEN_SWE_BOT_NAME,
resolve_triggering_user_identity,
)
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
from .utils.dashboard_links import dashboard_plan_url, dashboard_thread_url
from .utils.github_app import (
get_github_app_installation_token_with_expiry,
)
from .utils.github_proxy import record_proxy_token_expiry
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) * fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
from .utils.github_token import repo_cache_key
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
from .utils.model import (
fallback_model_id_for,
make_model,
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) * feat: tune reviewer for precision — web/wiki tools + recalibrated prompt Reviewer agent now has web_search, fetch_url, and http_request alongside the finding tools, so it can verify library semantics and consult the DeepWiki auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>) before flagging cross-file or architectural concerns. Prompt rewritten to push precision over recall: - explicit severity ladder pushing reviews toward bimodal high/low instead of defaulting to medium - ≤200-char description target (gold set averages ~186 chars; we were at ~436) - mandatory docs / wiki / code lookup before flagging concurrency, security, or perf — the three categories that dominated false positives - "do not flag" list covering compiler/linter-catchable nits, speculative claims without a concrete attacker/interleaving/scale, style preferences the codebase doesn't share, and test-quality nits on non-test diffs - smart file-selection guidance for large PRs (deprioritize generated / vendored / pure-rename hunks) Eval config switched to openai:gpt-5.5 + high reasoning effort for the next benchmark run. * trim prompt * subagent prompting * confidence ratings * added medium * enforce confidence threshold * . * reviewer: precision-tuned prompt + drop confidence gate Rewrites the reviewer system prompt around a defensibility bar (anchor + failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file list (style nits, speculation, scope-policing, same-bug fan-out), and a checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26% style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new prompt targets each class directly. Confidence is still recorded on every finding for post-hoc calibration but no longer gates publication — the audit showed the gate was a no-op (agent self-rated 65% of findings "high" regardless), and the prompt's defensibility bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD, the confidence_threshold kwarg on filter_findings_for_publish, the confidence_filtered score_mode, and the min_confidence kwarg on the eval target's _extract_comments — all dead once the gate is gone. Also removes the "informational" severity tier from the Severity enum, SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for FYI observations the dataset never rewards. * benchmax * adding google provider * slight steering * tuning * more tuning * fix * cleanup * reducing overfitting * Add per-repo review style profiles and inject them into the reviewer. Dashboard users can analyze historical PR review feedback per repository, edit the resulting style guide, and have it loaded from LangGraph Store at reviewer runtime (including Martian eval runs) keyed by owner/name. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix review style job errors leaking exception details to clients. Return generic dashboard messages while logging full stack traces server-side. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 11:35:00 -07:00
provider_model_kwargs,
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
)
from .utils.sandbox import create_sandbox
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
from .utils.sandbox_paths import aresolve_sandbox_work_dir
from .utils.tracing import AGENT_TRACING_PROJECT, traced_graph_factory
client = get_client()
SANDBOX_CREATING = "__creating__"
SANDBOX_CREATION_TIMEOUT = 180
SANDBOX_POLL_INTERVAL = 1.0
from .utils.sandbox_state import (
SANDBOX_BACKENDS,
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) * fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
get_bound_repo_from_metadata,
get_sandbox_id_from_metadata,
set_sandbox_backend,
unwrap_sandbox_backend,
)
async def _resolve_prompt_default_repo(configurable: dict[str, Any]) -> dict[str, str] | None:
repo_config = configurable.get("repo")
if isinstance(repo_config, dict):
owner = repo_config.get("owner")
name = repo_config.get("name")
if isinstance(owner, str) and isinstance(name, str):
return {"owner": owner, "name": name}
if configurable.get("repo_explicitly_none") is True:
return None
try:
return await get_team_default_repo()
except Exception:
logger.debug("Failed to load team default repo for prompt", exc_info=True)
return None
async def _resolve_repo_custom_instructions(
default_repo: dict[str, str] | None,
) -> str | None:
"""Load per-repo custom agent instructions for the resolved default repo."""
if not default_repo or not default_repo.get("owner") or not default_repo.get("name"):
return None
try:
from .dashboard.agent_instructions import get_repo_agent_instructions
return await get_repo_agent_instructions(default_repo["owner"], default_repo["name"])
except Exception:
logger.debug("Failed to load repo custom agent instructions", exc_info=True)
return None
async def _start_langsmith_sandbox_if_needed(sandbox_backend: SandboxBackendProtocol) -> None:
"""Start a LangSmith sandbox before operations that require it to be running."""
if os.getenv("SANDBOX_TYPE", "langsmith") != "langsmith":
return
current_backend = unwrap_sandbox_backend(sandbox_backend)
if not isinstance(current_backend, LangSmithSandbox):
return
sandbox = current_backend._sandbox # noqa: SLF001
status = await asyncio.to_thread(sandbox._client.get_sandbox_status, sandbox.name) # noqa: SLF001
status_name = getattr(status, "status", status)
status_name = getattr(status_name, "value", status_name)
status_text = str(status_name or "").lower()
if status_text in {"running", "ready"}:
return
logger.info(
"Starting LangSmith sandbox %s before proxy refresh (status=%s)",
current_backend.id,
status_text or "unknown",
)
await asyncio.to_thread(sandbox.start)
async def _resolve_proxy_token(github_proxy_token: str | None) -> tuple[str | None, str | None]:
"""Resolve the proxy token and its expiry.
An explicitly supplied token has no known expiry; otherwise we mint a fresh
GitHub App installation token and keep its ``expires_at`` so the proxy can
be refreshed before the (hard 1h) expiry.
"""
if github_proxy_token:
return github_proxy_token, None
return await get_github_app_installation_token_with_expiry()
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
async def _resolve_snapshot_id_for_repo(repo: dict[str, str] | None) -> str | None:
"""Resolve a repo's ready snapshot id; ``None`` falls back to the default.
Never raises: any failure resolves to ``None`` so sandbox creation falls
back to the configured ``DEFAULT_SANDBOX_SNAPSHOT_ID``.
"""
if not repo:
return None
try:
return await resolve_repo_snapshot_id(repo.get("owner"), repo.get("name"))
except Exception: # noqa: BLE001
logger.debug("Failed to resolve repo-scoped snapshot", exc_info=True)
return None
async def _create_sandbox_with_proxy(
github_proxy_token: str | None = None,
*,
thread_id: str | None = None,
github_proxy_repositories: Sequence[str] | None = None,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo: dict[str, str] | None = None,
) -> SandboxBackendProtocol:
"""Create a new sandbox with GitHub proxy auth configured."""
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
snapshot_id = await _resolve_snapshot_id_for_repo(repo)
sandbox_backend = await asyncio.to_thread(create_sandbox, snapshot_id=snapshot_id)
sandbox_type = os.getenv("SANDBOX_TYPE", "langsmith")
if sandbox_type == "langsmith":
token, expires_at = await _resolve_proxy_token(github_proxy_token)
if not token:
msg = "Cannot configure proxy: GitHub App installation token is unavailable"
logger.error(msg)
raise ValueError(msg)
await _start_langsmith_sandbox_if_needed(sandbox_backend)
await asyncio.to_thread(_configure_github_proxy, sandbox_backend.id, token)
record_proxy_token_expiry(thread_id, expires_at, repositories=github_proxy_repositories)
return sandbox_backend
async def _refresh_github_proxy(
sandbox_backend: SandboxBackendProtocol,
github_proxy_token: str | None = None,
*,
thread_id: str | None = None,
github_proxy_repositories: Sequence[str] | None = None,
) -> None:
"""Refresh GitHub proxy credentials for reused LangSmith sandboxes."""
if os.getenv("SANDBOX_TYPE", "langsmith") != "langsmith":
return
token, expires_at = await _resolve_proxy_token(github_proxy_token)
if not token:
logger.warning(
"Skipping GitHub proxy refresh for sandbox %s: installation token unavailable",
sandbox_backend.id,
)
return
current_backend = unwrap_sandbox_backend(sandbox_backend)
await _start_langsmith_sandbox_if_needed(current_backend)
await asyncio.to_thread(_configure_github_proxy, current_backend.id, token)
record_proxy_token_expiry(thread_id, expires_at, repositories=github_proxy_repositories)
async def _refresh_github_proxy_or_recreate(
sandbox_backend: SandboxBackendProtocol,
thread_id: str,
github_proxy_token: str | None = None,
github_proxy_repositories: Sequence[str] | None = None,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo: dict[str, str] | None = None,
) -> SandboxBackendProtocol:
"""Refresh proxy credentials, recreating stale LangSmith sandboxes on failure."""
try:
await _refresh_github_proxy(
sandbox_backend,
github_proxy_token,
thread_id=thread_id,
github_proxy_repositories=github_proxy_repositories,
)
except Exception: # noqa: BLE001
logger.warning(
"Failed to refresh GitHub proxy for sandbox %s on thread %s, recreating sandbox",
sandbox_backend.id,
thread_id,
exc_info=True,
)
return await _recreate_sandbox(
thread_id,
github_proxy_token=github_proxy_token,
github_proxy_repositories=github_proxy_repositories,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo=repo,
)
return sandbox_backend
async def _configure_git_identity(sandbox_backend: SandboxBackendProtocol) -> None:
await asyncio.to_thread(
sandbox_backend.execute,
f"git config --global user.name '{OPEN_SWE_BOT_NAME}' && "
f"git config --global user.email '{OPEN_SWE_BOT_EMAIL}'",
)
async def _recreate_sandbox(
thread_id: str,
*,
github_proxy_token: str | None = None,
github_proxy_repositories: Sequence[str] | None = None,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo: dict[str, str] | None = None,
) -> SandboxBackendProtocol:
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
"""Recreate a sandbox after a connection failure.
Sets the SANDBOX_CREATING sentinel and creates a fresh sandbox
(with proxy auth configured), swapping the per-thread proxy target.
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
The agent is responsible for cloning repos via tools.
"""
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
await client.threads.update(thread_id=thread_id, metadata=_creating_metadata())
try:
sandbox_backend = set_sandbox_backend(
thread_id,
await _create_sandbox_with_proxy(
github_proxy_token,
thread_id=thread_id,
github_proxy_repositories=github_proxy_repositories,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo=repo,
),
)
except Exception:
logger.exception("Failed to recreate sandbox after connection failure")
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
await client.threads.update(thread_id=thread_id, metadata=_RESET_METADATA)
raise
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
return sandbox_backend
async def check_or_recreate_sandbox(
sandbox_backend: SandboxBackendProtocol,
thread_id: str,
github_proxy_token: str | None = None,
github_proxy_repositories: Sequence[str] | None = None,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo: dict[str, str] | None = None,
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
) -> SandboxBackendProtocol:
"""Check if a cached sandbox is reachable; recreate it if not.
Pings the sandbox with a lightweight command. If the sandbox is
unreachable (SandboxClientError), it is torn down and a fresh one
is created via _recreate_sandbox.
Returns the original backend if healthy, or a new one if recreated.
"""
try:
await asyncio.to_thread(sandbox_backend.execute, "echo ok")
except SandboxClientError:
logger.warning(
"Cached sandbox is no longer reachable for thread %s, recreating",
thread_id,
)
sandbox_backend = await _recreate_sandbox(
thread_id,
github_proxy_token=github_proxy_token,
github_proxy_repositories=github_proxy_repositories,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo=repo,
)
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
return sandbox_backend
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
def _creating_metadata() -> dict[str, Any]:
"""Metadata that claims the cross-process creation lock with a timestamp."""
return {"sandbox_id": SANDBOX_CREATING, "sandbox_creating_at": time.time()}
_RESET_METADATA: dict[str, Any] = {"sandbox_id": None, "sandbox_creating_at": None}
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
async def _resolve_creating_sentinel(thread_id: str) -> str | None:
"""Resolve a ``__creating__`` sentinel seen with no cached backend.
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
The sentinel is a cross-process lock: another worker may still be creating
the sandbox. Poll live thread metadata until it resolves to a real id. Only
when the sentinel is older than ``SANDBOX_CREATION_TIMEOUT`` (e.g. the
creating worker was restarted) is it treated as stale: metadata is reset and
``None`` is returned so the caller creates a fresh sandbox. A sentinel with
no timestamp (written before this field existed) is also treated as stale.
"""
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
while True:
thread = await client.threads.get(thread_id)
metadata = thread.get("metadata", {}) if isinstance(thread, dict) else {}
sandbox_id = metadata.get("sandbox_id") if isinstance(metadata, dict) else None
if sandbox_id != SANDBOX_CREATING:
return sandbox_id if isinstance(sandbox_id, str) else None
creating_at = metadata.get("sandbox_creating_at") if isinstance(metadata, dict) else None
age = time.time() - creating_at if isinstance(creating_at, (int, float)) else None
if age is None or age > SANDBOX_CREATION_TIMEOUT:
logger.warning(
"Resetting stale SANDBOX_CREATING for thread %s (age=%s)", thread_id, age
)
await client.threads.update(thread_id=thread_id, metadata=_RESET_METADATA)
return None
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
await asyncio.sleep(SANDBOX_POLL_INTERVAL)
def graph_loaded_for_execution(config: RunnableConfig) -> bool:
"""Check if the graph is loaded for actual execution vs introspection."""
return (
config["configurable"].get("__is_for_execution__", False)
if "configurable" in config
else False
)
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) * fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
class SandboxRepoMismatchError(RuntimeError):
"""Raised when a thread_id is presented for a repo it is not bound to.
A thread is bound to exactly one repo. A different repo presenting a
colliding thread_id (e.g. an attacker-named branch whose first UUID matches
another thread) must never reuse this thread's sandbox or token.
"""
def __init__(self, thread_id: str, bound_repo: str, current_repo: str) -> None:
self.thread_id = thread_id
self.bound_repo = bound_repo
self.current_repo = current_repo
super().__init__(
f"Thread {thread_id} is bound to repo {bound_repo}, "
f"refusing to serve sandbox for {current_repo}"
)
async def ensure_sandbox_for_thread(
thread_id: str,
*,
github_proxy_token: str | None = None,
github_proxy_repositories: Sequence[str] | None = None,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo: dict[str, str] | None = None,
) -> SandboxBackendProtocol:
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
"""Get-or-create a healthy sandbox bound to ``thread_id``.
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
Implements the four-state lifecycle described in AGENTS.md:
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
1. Cached in memory → ping; recreate on ``SandboxClientError``.
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
2. Metadata says ``__creating__`` and no cache → wait for the creating
worker; only reset if the sentinel is proven stale (timestamp/timeout).
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
3. No sandbox at all → create one and persist the id.
4. Metadata has an id but no cache → reconnect; recreate on failure.
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
For LangSmith sandboxes, also refreshes the GitHub App proxy auth. When
``repo`` has a ``ready`` repo-scoped snapshot, newly created sandboxes boot
from it; otherwise the configured ``DEFAULT_SANDBOX_SNAPSHOT_ID`` is used.
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
Persists the resulting ``sandbox_id`` to thread metadata, and on the
first creation/reconnect for this thread initializes git identity.
"""
sandbox_backend = SANDBOX_BACKENDS.get(thread_id)
sandbox_id = await get_sandbox_id_from_metadata(thread_id)
if sandbox_id == SANDBOX_CREATING and not sandbox_backend:
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
logger.info("Sandbox creation in progress for thread %s, waiting...", thread_id)
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
sandbox_id = await _resolve_creating_sentinel(thread_id)
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) * fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
# Repo-binding guard (TID-COLLIDE-01): a sandbox is never served to a repo
# unless its binding is known and matches.
current_repo = repo_cache_key(repo)
bound_repo = await get_bound_repo_from_metadata(thread_id)
proxy_bound = getattr(sandbox_backend, "bound_repo", None)
effective_bound = bound_repo or (proxy_bound if isinstance(proxy_bound, str) else None)
if current_repo and effective_bound and effective_bound != current_repo:
# Known binding that does not match the current repo: refuse outright so a
# colliding thread_id from a different repo cannot reuse/clobber it.
logger.error(
"Repo mismatch for thread %s: bound=%s current=%s; refusing sandbox reuse",
thread_id,
effective_bound,
current_repo,
)
raise SandboxRepoMismatchError(thread_id, effective_bound, current_repo)
if (
current_repo
and not effective_bound
and sandbox_backend is None
and isinstance(sandbox_id, str)
and sandbox_id not in (None, SANDBOX_CREATING)
):
# Fail CLOSED for unbound-legacy threads (migration window): a thread with a
# persisted sandbox_id but no in-memory cache and no recorded bound_repo
# cannot be confirmed to belong to the current repo, so never
# reconnect-and-serve it. Drop the stale id and recreate a fresh sandbox
# bound to this repo below.
logger.error(
"reconnect-with-missing-binding for thread %s: persisted sandbox %s has no "
"bound_repo; refusing reuse and recreating for repo %s",
thread_id,
sandbox_id,
current_repo,
)
sandbox_id = None
if sandbox_backend:
logger.info("Using cached sandbox backend for thread %s", thread_id)
original_sandbox_id = sandbox_backend.id
sandbox_backend = await check_or_recreate_sandbox(
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
sandbox_backend, thread_id, github_proxy_token, github_proxy_repositories, repo
)
if sandbox_backend.id == original_sandbox_id:
sandbox_backend = await _refresh_github_proxy_or_recreate(
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
sandbox_backend, thread_id, github_proxy_token, github_proxy_repositories, repo
)
elif sandbox_id is None:
logger.info("Creating new sandbox for thread %s", thread_id)
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
await client.threads.update(thread_id=thread_id, metadata=_creating_metadata())
try:
sandbox_backend = await _create_sandbox_with_proxy(
github_proxy_token,
thread_id=thread_id,
github_proxy_repositories=github_proxy_repositories,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo=repo,
)
logger.info("Sandbox created: %s", sandbox_backend.id)
except Exception:
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
logger.exception("Failed to create sandbox")
try:
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
await client.threads.update(thread_id=thread_id, metadata=_RESET_METADATA)
except Exception:
logger.exception("Failed to reset sandbox_id metadata")
raise
else:
logger.info("Connecting to existing sandbox %s", sandbox_id)
created_replacement_sandbox = False
try:
sandbox_backend = await asyncio.to_thread(create_sandbox, sandbox_id)
except Exception:
logger.warning("Failed to connect to existing sandbox %s, creating new one", sandbox_id)
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
await client.threads.update(thread_id=thread_id, metadata=_creating_metadata())
try:
sandbox_backend = await _create_sandbox_with_proxy(
github_proxy_token,
thread_id=thread_id,
github_proxy_repositories=github_proxy_repositories,
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
repo=repo,
)
created_replacement_sandbox = True
except Exception:
logger.exception("Failed to create replacement sandbox")
feat: outcomes dataset + bootstrap/continual split via skills (#1365) * fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
await client.threads.update(thread_id=thread_id, metadata=_RESET_METADATA)
raise
if not created_replacement_sandbox:
original_sandbox_id = sandbox_backend.id
sandbox_backend = await check_or_recreate_sandbox(
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
sandbox_backend, thread_id, github_proxy_token, github_proxy_repositories, repo
)
if sandbox_backend.id == original_sandbox_id:
sandbox_backend = await _refresh_github_proxy_or_recreate(
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
sandbox_backend, thread_id, github_proxy_token, github_proxy_repositories, repo
)
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) * fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
sandbox_backend = set_sandbox_backend(thread_id, sandbox_backend, repo=current_repo)
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) * fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
metadata_update: dict[str, Any] = {}
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
if sandbox_id != sandbox_backend.id:
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) * fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
metadata_update["sandbox_id"] = sandbox_backend.id
if current_repo and bound_repo != current_repo:
metadata_update["bound_repo"] = current_repo
if metadata_update:
await client.threads.update(thread_id=thread_id, metadata=metadata_update)
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
# Re-apply git identity every run: cached/reconnected sandboxes may have
# lost their `--global` config (or had it overwritten), and Vercel preview
# deploys reject commits whose author email can't be resolved to a GitHub
# account.
await _configure_git_identity(sandbox_backend)
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
return sandbox_backend
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
DEFAULT_LLM_MODEL_ID = DEFAULT_MODEL_ID
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
DEFAULT_LLM_MAX_TOKENS = 64_000
DEFAULT_RECURSION_LIMIT = 9_999
MODEL_CALL_RECURSION_LIMIT = 5_000 # ~half the recursion limit to account for tool calls
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
# Mutating tools hidden from the model while plan mode is active so it can only
# research and propose a plan. `execute` stays available; plan-mode shell
# discipline (no mutating commands) is instructed via the system prompt rather
# than enforced. `http_request` is excluded because it can POST/PUT/PATCH/DELETE
# to external services — read-only web research goes through `web_search` /
# `fetch_url`. `task` is excluded because the general-purpose subagent is built
# with its own filesystem/PR/Linear tools and does not inherit this exclusion, so
# delegating to it would bypass the read-only intent.
PLAN_MODE_EXCLUDED_TOOLS: frozenset[str] = frozenset(
{
"write_file",
"edit_file",
"task",
"http_request",
"open_pull_request",
"request_pr_review",
"linear_create_issue",
"linear_update_issue",
"linear_delete_issue",
}
)
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
def _general_purpose_subagent(model: BaseChatModel) -> SubAgent:
return {
"name": GENERAL_PURPOSE_SUBAGENT["name"],
"description": GENERAL_PURPOSE_SUBAGENT["description"],
"system_prompt": GENERAL_PURPOSE_SUBAGENT["system_prompt"],
"model": model,
}
def _get_cached_sandbox_backend(thread_id: str) -> SandboxBackendProtocol:
sandbox_backend = SANDBOX_BACKENDS.get(thread_id)
if sandbox_backend is None:
raise RuntimeError(f"No sandbox backend cached for thread {thread_id}")
return sandbox_backend
async def _observability_authorized(config: RunnableConfig, profile_login: str | None) -> bool:
"""Whether the triggering user may use the team observability tools.
Gates on admin / explicitly-authorized emails so prompt-injected runs from
untrusted contributors cannot reach the team's Datadog/LangSmith data.
"""
configurable = (config or {}).get("configurable") or {}
slack_thread = configurable.get("slack_thread") or {}
config_login = configurable.get("github_login")
candidate_login = profile_login or (config_login if isinstance(config_login, str) else None)
candidate_emails = [
configurable.get("user_email"),
slack_thread.get("triggering_user_email"),
]
if any(is_observability_authorized(email, login=candidate_login) for email in candidate_emails):
return True
return is_observability_authorized(
await email_for_login(candidate_login), login=candidate_login
)
async def _load_observability_tools(authorized: bool) -> list[Any]:
"""Datadog (MCP) + LangSmith read tools when the team has connected them.
Credentials live server-side in team settings; the sandbox never holds them.
Only loaded for authorized (admin / allow-listed) triggering users so an
untrusted run cannot exfiltrate team observability data. Failures degrade to
no tools so the agent still starts.
"""
if not authorized:
return []
try:
datadog_tools, langsmith_tools = await asyncio.gather(
load_datadog_tools(),
load_langsmith_tools(),
)
except Exception:
logger.warning("Failed to load observability tools", exc_info=True)
return []
return [*datadog_tools, *langsmith_tools]
async def _load_corridor_mcp_tools() -> list[Any]:
"""Corridor MCP tools when the deployment environment has configured them."""
try:
return await load_corridor_tools()
except Exception:
logger.warning("Failed to load Corridor MCP tools", exc_info=True)
return []
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
async def get_agent(config: RunnableConfig) -> Pregel:
"""Get or create an agent with a sandbox for the given thread."""
thread_id = config["configurable"].get("thread_id", None)
config["recursion_limit"] = DEFAULT_RECURSION_LIMIT
if thread_id is None or not graph_loaded_for_execution(config):
logger.info("No thread_id or not for execution, returning agent without sandbox")
return create_deep_agent(
system_prompt="",
tools=[],
).with_config(config)
github_token, _expires_at = await resolve_github_token(config, thread_id)
profile_login = resolve_github_login(config)
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
configurable = (config or {}).get("configurable") or {}
prompt_default_repo = await _resolve_prompt_default_repo(configurable)
feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60) * feat: default Slack/dashboard/schedule PRs + commits to the app identity (#57) Slack/dashboard/schedule runs now author PRs and run git/gh operations as the GitHub App seahaven-openswe[bot] by default (matching GitHub-issue runs), so the self-review 422 is impossible by construction rather than guarded in the prompt. A profile flag author_prs_as_user restores per-user attribution. - open_pull_request._resolve_pr_author_token + auth.resolve_github_token: default to the installation token for these sources; per-user only when opted in. - authorship: commit identity -> seahaven-openswe[bot] (numeric noreply; accepted Vercel-resolution risk, documented inline). - self-trigger safety: INTERNAL_BOT_LOGINS + webapp/reviewer_reconcile/reply markers recognize seahaven-openswe[bot] (bot-authored events are now ours). Supersedes the prompt-only guard in #58. * fix: author commits as the app bot in the default path (SH-IDSPLIT-01) Security review found the commit identity was NOT actually unified to the bot: resolve_triggering_user_identity got a 403 from the installation token and fell back to configurable['github_login'], so commits were still authored as the triggering user (commit=user, push+PR=bot — a three-way split that missed the stated goal). Now gate the triggering-user identity resolution on the same default-bot decision as the token: slack/dashboard/schedule default to the app bot identity unless author_prs_as_user is set. * docs(security): record AUTHZ-SLACK-BOT-DEFAULT-001 as an accepted residual (#59) Single-user deployment; bounded by App-on-pilot + ALLOWED_GITHUB_REPOS lock. Revisit (add a per-user gate) before expanding users or the App installation.
2026-06-29 14:22:33 -04:00
# Commit identity must follow the SAME default-bot decision as the token
# (SH-IDSPLIT-01): by default slack/dashboard/schedule runs author commits as the
# app bot, so resolve the triggering USER's git identity ONLY when authoring as the
# user (the author_prs_as_user opt-in, or a non-default source). Otherwise leave it
# None so construct_system_prompt sets the bot identity (OPEN_SWE_BOT_NAME/EMAIL) and
# commits don't get mis-attributed to a human who didn't write them.
if configurable.get("source") in ("slack", "dashboard", "schedule"):
_author_as_user = bool(
isinstance(profile_login, str)
and profile_login.strip()
and profile_author_prs_as_user(await load_profile(profile_login.strip()))
)
else:
_author_as_user = True
async def _no_triggering_identity() -> Any:
return None
if _author_as_user:
triggering_user_identity_task = asyncio.create_task(
asyncio.to_thread(resolve_triggering_user_identity, config, github_token)
)
else:
triggering_user_identity_task = asyncio.create_task(_no_triggering_identity())
feat: repo-scoped dynamic sandbox snapshots (#1595) * feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
sandbox_task = asyncio.create_task(
ensure_sandbox_for_thread(thread_id, repo=prompt_default_repo)
)
team_defaults_task = asyncio.create_task(get_team_default_model_pair("agent"))
profile_task = asyncio.create_task(load_profile(profile_login)) if profile_login else None
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) * fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
try:
triggering_user_identity, sandbox_backend, team_defaults = await asyncio.gather(
triggering_user_identity_task,
sandbox_task,
team_defaults_task,
)
except SandboxRepoMismatchError as exc:
# Repo-binding refusal at the run boundary: log for alarming and surface the
# already-sanitized terminal error (no sandbox/token internals) to the caller,
# rather than letting an opaque deep-stack exception crash-loop the worker.
logger.error("Refusing agent run for thread %s: %s", thread_id, exc)
for pending in (triggering_user_identity_task, team_defaults_task, profile_task):
if pending is not None and not pending.done():
pending.cancel()
raise RuntimeError(str(exc)) from exc
profile = await profile_task if profile_task is not None else None
feat: add reviewer graph + eval target wiring (#1241) * feat: add reviewer graph + eval target wiring - New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json alongside the main `agent` graph. Reuses the same sandbox lifecycle, GH proxy auth, and middleware primitives from `agent.server`, but with a narrower tool set, a reviewer-specific system prompt, no commit/push, and the `task` (subagent) tool stripped via `_ToolExclusionMiddleware` so review stays in one context. - New `github_comment` tool: agents call it once per issue with `(file, line, body, severity)` and the eval scores those calls against golden comments. - `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally *not* on the reviewer's stack — that middleware exists to enforce the main agent's "always finalize via Slack/Linear/PR" contract, which the reviewer doesn't have. The main agent's behavior is unchanged. - `evals/reviewer/target.py`: send PR info as a user message, extract every `github_comment` tool call (multiple expected per review) into the run output. - `evals/reviewer/judge.py`: per-example evaluator now returns a list of metrics under `{"results": [...]}` so LangSmith averages each numeric key (f1/precision/recall/tp/fp/fn) across the experiment in the UI. Dropped the broken `aggregate_pr` summary evaluator that reached for an attribute that doesn't exist on `RunTree`. - `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via `client.list_examples(limit=N)` since `aevaluate` doesn't accept `max_examples`. - Makefile: `dev` and `run` targets now use `uv run` so they work without an activated venv. * resolve comments --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
del github_token
linear_issue = config["configurable"].get("linear_issue", {})
linear_project_id = linear_issue.get("linear_project_id", "")
linear_issue_number = linear_issue.get("linear_issue_number", "")
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
work_dir = await aresolve_sandbox_work_dir(sandbox_backend)
def backend_factory(_runtime: object, _thread_id: str = thread_id) -> SandboxBackendProtocol:
return _get_cached_sandbox_backend(_thread_id)
(model_id, profile_effort), (subagent_model_id, subagent_effort) = team_defaults
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
logger.info("Using team default agent model: model=%s effort=%s", model_id, profile_effort)
logger.info(
"Using team default agent subagent model: model=%s effort=%s",
subagent_model_id,
subagent_effort,
)
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
if profile_login and profile:
overridden_model, overridden_effort = normalize_profile_overrides(profile)
if overridden_model:
logger.info(
"Applying dashboard profile override for %s: model=%s effort=%s",
profile_login,
overridden_model,
overridden_effort,
)
model_id = overridden_model
profile_effort = overridden_effort
subagent_model_id = overridden_model
subagent_effort = overridden_effort
overridden_subagent_model, overridden_subagent_effort = (
normalize_profile_subagent_overrides(profile)
)
if overridden_subagent_model:
logger.info(
"Applying dashboard profile subagent override for %s: model=%s effort=%s",
profile_login,
overridden_subagent_model,
overridden_subagent_effort,
)
subagent_model_id = overridden_subagent_model
subagent_effort = overridden_subagent_effort
feat: open-swe dashboard for per-user profile config (#1302) * feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it.
2026-05-15 11:23:53 -07:00
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
per_thread_model = configurable.get("agent_model_id")
per_thread_effort = configurable.get("agent_effort")
if (
isinstance(per_thread_model, str)
and per_thread_model in SUPPORTED_MODEL_IDS
and isinstance(per_thread_effort, str)
and model_supports_effort(per_thread_model, per_thread_effort)
):
logger.info(
"Applying per-thread model override: model=%s effort=%s",
per_thread_model,
per_thread_effort,
)
model_id = per_thread_model
profile_effort = per_thread_effort
subagent_model_id = per_thread_model
subagent_effort = per_thread_effort
feat: add Agents chat UI for cloud threads (#1323) * feat(ui): add Agents chat UI ported from open-swe-app Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dashboard): wire Agents UI to LangGraph thread APIs Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dashboard): single agent reply per turn in Agents UI Use UUID thread IDs LangGraph accepts, skip confirming_completion for dashboard threads, and merge adapter agent messages so duplicate bubbles do not render. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): polish Agents UI with floating prompt and layout cleanup Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar from open-swe-app, and refine chat layout so messages scroll behind the input. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): patch deepagents reducer for None messages on checkpoint replay LangGraph thread state could 500 when cancelled runs left messages as None. Apply the reducer guard before graph import, fall back to metadata in the dashboard API, and adjust Agents prompt bar layout. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(ui): unify sidebar user menu and clean up Agents UI navigation Extract SidebarUserMenu so the dashboard and Agents sidebars render the same profile button, drop the redundant Agents nav row in favor of the existing Back to Agents link, add the open-swe logo header to the Agents sidebar, flatten the New Agent button, and cap the home screen run list to keep the prompt input in view. * feat(ui): resizable/collapsible sidebar shared across dashboard and Agents Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars share a persisted width (default 260px, drag to resize, 200-420 range) and a collapse toggle that hides the panel and surfaces a floating reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on- hover thread delete control in the Agents sidebar. * feat(ui): instant user message and busy indicator on Agents transition Stash submitted prompts in sessionStorage, pre-populate the new thread detail cache, and merge pending prompts into the rendered message list so the Agents page renders the user bubble plus the existing thinking spinner immediately instead of flashing a skeleton and "Agent is starting" while the run boots. * feat(ui): token-stream agent replies in the Agents thread view Opt the LangGraph runs into messages-tuple streaming and forward those events through the existing SSE channel. The frontend now applies AIMessageChunk deltas directly to the cached thread (cancelling any in-flight refetch first so optimistic tokens are not clobbered) and keeps positional pending prompts so the user bubble stays in the right place while the agent streams its reply. * fix(dashboard): await threads.join_stream before iterating threads.join_stream is async def returning an AsyncIterator, so it must be awaited before async for. The SSE endpoint was raising TypeError: 'async for' requires an object with __aiter__ method, got coroutine on every connection. * fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns Setting stream_mode=["values","messages-tuple","updates"] on runs.create forces langchain_anthropic into streaming, and on the second model call (after tool execution) its serialized thinking blocks come back malformed, so Anthropic rejects the request with 'messages.1.content.0.thinking.thinking: Field required'. Revert to the default stream_mode so claude-opus thinking + tool use runs to completion. The frontend keeps the messages-event handler in place as a no-op fallback for when streaming is re-enabled. * feat(agents): per-thread model picker wired through to the run Add optional model_id/effort to the create-thread and send-message request bodies, forward them as agent_model_id/agent_effort in the LangGraph run configurable, and record the resolved choice in thread metadata so the UI can show the model the run is actually using. get_agent now picks the per-thread override last (highest priority over team default + profile override) and falls back gracefully when it is absent or unsupported. The frontend prompt bar becomes a controlled component fed by a shared useModelOptions hook (options + profile -> defaultSelection). AgentsHome seeds the picker from the user's profile default; the thread view seeds from the thread's recorded model/effort and lets each follow-up retarget the run. * refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar Drop the absolute-positioned send button, restore the original px-4 py-3.5 min-h-[106px] flex-col container, and move the model picker into a mt-auto pt-2 footer row so the placeholder text and the model selector share the same horizontal padding. * chore: fix lint/format CI failures Remove unused imports and reformat two files flagged by ruff. * fix(tests): stop messages-reducer patch tests from polluting the suite Restore agent modules after reducer patch tests and import LangSmithSandbox from agent.server in proxy refresh tests so isinstance checks stay valid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 11:15:59 -07:00
always_create_prs = profile_create_prs(profile)
if always_create_prs:
logger.info("Always Create PRs enabled by profile for %s", profile_login)
feat: restructure Open SWE Review tab + wire create_prs (#1319) * feat(dashboard): restructure Open SWE Review tab + wire create_prs Restructures the dashboard around two related changes the reviewer settings have been asking for: - Wire profile.create_prs. Defaults to true (opt-out); when off the system prompt gets a `Pull Request Policy Override` section telling the agent to push the branch and notify with the branch URL instead of opening a PR. Removes the noop Slack Notifications / Allow Artifacts / First Name / Last Name controls and their schema fields. - Repositories opt-in for Open SWE Review. New per-team enabled list stored in the LangGraph Store (`["enabled_review_repos"]`). Every reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review` which AND-combines the existing env allowlist with the dashboard list. Default is empty (opt-in) — admins enable repos per-installation from the new Repositories page nested under Open SWE Review. - Open SWE Review tab now mirrors the Cursor "rules" pattern: main page shows installation rows + a Rules entry; both drill into nested pages (/review/repositories/$owner and /review/styles) with a back link. - Adds the new logo/favicon assets shipped from sidebar + html head. Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults `is_review_repo_enabled` to True for existing allowlist tests. * fix(dashboard): make main content scroll independently of the sidebar Outer flex container was min-h-svh, so it grew with main's content and the whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden so the sidebar stays put and only <main> scrolls. * fix(dashboard): make disabled repo toggles obviously disabled Switch's disabled state used opacity-50 against a muted background, so the not-admin state looked nearly identical to the off state. Bump to opacity-40 + grayscale, and wrap each repo toggle in a span carrying a native hover tooltip explaining why it's disabled. * fix(switch): handle base-ui's data-disabled state base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute) when disabled, so Tailwind's disabled: variant never matches and the button keeps its cursor-pointer + clickable look. Mirror the styling under the data-[disabled] variant and add pointer-events-none so the disabled state is both visible and actually unclickable. * feat(dashboard): paginate per-installation repository list 20 repos per page with Prev / page X of Y / Next controls at the bottom. Pager only renders when there are more than 20 repos. Page resets to 0 when navigating between installations. * feat(dashboard): global default model selectors for Agent + Reviewer Adds team-wide default model + reasoning effort for both agents in the Admin tab so operators can switch models without redeploying. Resolution chain: Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable Team defaults live in team_settings and are validated against the SUPPORTED_MODELS allowlist + the model's supported reasoning efforts. 'Inherit from env' clears the override and falls back to LLM_MODEL_ID. * refactor(models): drop LLM_MODEL_ID env in favour of the team default The team default is now the single source of truth for the runtime model choice; per-user (agent) and per-call configurable (reviewer) selections still win on top. When no admin has touched the team default, it surfaces the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the admin UI's dropdown is always pre-populated with a sensible value. The Admin UI loses the 'Inherit from env' option since there is no longer an env layer to inherit from. * chore(models): set hardcoded fallback to gpt-5.5 medium Decouple the team-default boot value (gpt-5.5 / medium) from each model's ProfileForm-suggested default_effort so we can change one without nudging the other. The Opus xhigh default for new user profiles is unchanged. * feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings - Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new description copy that matches the screenshot. Legacy stored values fall back to 'every_push' on read so the UI never shows an unknown selection. - Add a 'Coming soon' badge + greyed-out + disabled state on the controls that don't have runtime consumers yet: Trigger Mode, Autofix Mode, Autofix Severity Threshold, and Automatically fix CI failures. SettingsRow grew a comingSoon prop to keep this consistent. - My Settings drops the noop PR Preferences section and adds a Sign Out button. preferred_pr_destination is removed from the profile schema; old records get the field popped on next write. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) * feat: tune reviewer for precision — web/wiki tools + recalibrated prompt Reviewer agent now has web_search, fetch_url, and http_request alongside the finding tools, so it can verify library semantics and consult the DeepWiki auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>) before flagging cross-file or architectural concerns. Prompt rewritten to push precision over recall: - explicit severity ladder pushing reviews toward bimodal high/low instead of defaulting to medium - ≤200-char description target (gold set averages ~186 chars; we were at ~436) - mandatory docs / wiki / code lookup before flagging concurrency, security, or perf — the three categories that dominated false positives - "do not flag" list covering compiler/linter-catchable nits, speculative claims without a concrete attacker/interleaving/scale, style preferences the codebase doesn't share, and test-quality nits on non-test diffs - smart file-selection guidance for large PRs (deprioritize generated / vendored / pure-rename hunks) Eval config switched to openai:gpt-5.5 + high reasoning effort for the next benchmark run. * trim prompt * subagent prompting * confidence ratings * added medium * enforce confidence threshold * . * reviewer: precision-tuned prompt + drop confidence gate Rewrites the reviewer system prompt around a defensibility bar (anchor + failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file list (style nits, speculation, scope-policing, same-bug fan-out), and a checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26% style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new prompt targets each class directly. Confidence is still recorded on every finding for post-hoc calibration but no longer gates publication — the audit showed the gate was a no-op (agent self-rated 65% of findings "high" regardless), and the prompt's defensibility bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD, the confidence_threshold kwarg on filter_findings_for_publish, the confidence_filtered score_mode, and the min_confidence kwarg on the eval target's _extract_comments — all dead once the gate is gone. Also removes the "informational" severity tier from the Severity enum, SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for FYI observations the dataset never rewards. * benchmax * adding google provider * slight steering * tuning * more tuning * fix * cleanup * reducing overfitting * Add per-repo review style profiles and inject them into the reviewer. Dashboard users can analyze historical PR review feedback per repository, edit the resulting style guide, and have it loaded from LangGraph Store at reviewer runtime (including Martian eval runs) keyed by owner/name. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix review style job errors leaking exception details to clients. Return generic dashboard messages while logging full stack traces server-side. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 11:35:00 -07:00
model_kwargs = provider_model_kwargs(
model_id,
profile_effort,
max_tokens=DEFAULT_LLM_MAX_TOKENS,
)
subagent_model_kwargs = provider_model_kwargs(
subagent_model_id,
subagent_effort,
max_tokens=DEFAULT_LLM_MAX_TOKENS,
)
fallback_model_id = os.environ.get("LLM_FALLBACK_MODEL_ID") or fallback_model_id_for(model_id)
fallback_middleware: list[Any] = []
if fallback_model_id and fallback_model_id != model_id:
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) * feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude) Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash) entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay freely selectable for the agent and reviewer graphs and via team/profile defaults. - pyproject: add langchain-aws (ChatBedrockConverse + boto3) - options.py: Bedrock Claude entry + default; remove openai/google entries - model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget), region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/ FIREWORKS_API_KEY local-dev validation - server.py: provider-aware fallback kwargs build - sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks - model_fallback: treat transient botocore ClientError codes as fallback-worthy - eval_jobs: repoint hardcoded eval model id to Bedrock Claude - tests: repoint dropped model ids; drop obsolete google test module * fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8 The handoff spec wired Bedrock Converse thinking as {type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a ValidationException: thinking.type "enabled" is not supported; it requires thinking.type "adaptive" plus output_config.effort. Verified by live invoke against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1): the enabled+budget shape 400s, adaptive+effort returns normally. Map profile effort to additional_model_request_fields: {thinking: {type: adaptive, display: summarized}, output_config: {effort: <low|medium|high|xhigh|max>}} reusing anthropic_thinking_for/anthropic_effort_for. Update the two subagent-model tests asserting the old shape. * fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids Model selection is store-driven, so seed_store.sh's team_settings/default seed is what runs in prod. It still seeded the removed providers, which would fail at runtime after the migration: - agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8 - reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8 (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer) - fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY -> FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would otherwise fail-fast at boot) - docs (DEPLOYMENT/ROTATION/put-config) updated to match. Surfaced by the cross-family review + verified against deploy/. * fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip From /sh-security-review (all confirmed-low): - model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches validate_local_dev_llm_config) so the validated region is the one actually used. - model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the error code only, so the role ARN + account id in the raw botocore message never reach logs or the user channel (CWE-209). - sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks (Converse emits reasoning_content, not thinking) so the middleware is not a no-op on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.) * deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids Deployment-readiness for the Bedrock migration (PR #62): - instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in each routed region (us-east-1/2, us-west-2). The model runs in the server process on the box, so the EC2 instance role is the principal. Simulator-verified (allowed for opus-4-8, implicitDeny for other models) and synth-verified. Passed the mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed). - config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides seed_store.sh's default via pick precedence, so the seed-script fix alone was insufficient — both sources now point at the supported Bedrock id. - infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai: ids to the Bedrock id (config.toml's model_id was an active, now-broken value). AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there. * chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed) Those three providers were dropped in the Bedrock/Fireworks migration and their keys revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY) were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not recreate the shells, and from fetch-config's mirror array so boot stops requesting them: - config-store.ts SECRET_VARS + descriptions (28 -> 25 shells) - fetch-config.sh SECRET_VARS array (kept in lockstep) - put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional (eval judge only — Bedrock builder/reviewer auth via the host IAM role). REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00
fallback_kwargs = provider_model_kwargs(
fallback_model_id, None, max_tokens=DEFAULT_LLM_MAX_TOKENS
)
fallback_middleware.append(
ModelFallbackMiddleware(make_model(fallback_model_id, **fallback_kwargs))
)
logger.info("Configured model fallback %s -> %s", model_id, fallback_model_id)
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
# Plan mode is entered only when the model decides to (the `enter_plan_mode`
# tool sets it in run state). The configurable value just carries that
# decision across a thread's messages and the approve/reject follow-ups; a
# fresh run with nothing set starts out of plan mode.
plan_mode = configurable.get("plan_mode") is True
if plan_mode:
logger.info("Plan mode enabled for thread %s", thread_id)
# Installed unconditionally and state-aware: it also restricts tools after a
# mid-run `enter_plan_mode` call, not just when plan mode is set up front.
plan_mode_middleware: list[Any] = [
PlanModeMiddleware(excluded=PLAN_MODE_EXCLUDED_TOOLS, initial=plan_mode)
]
source = (
configurable.get("source") if isinstance(configurable.get("source"), str) else "dashboard"
)
user_email = configurable.get("user_email")
user_email = user_email if isinstance(user_email, str) else ""
try:
await client.threads.update(
thread_id=thread_id,
metadata={
"agent_kind": "agent",
"model": model_id,
"effort": profile_effort,
"source": source,
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
"plan_mode": plan_mode,
},
)
await record_agent_thread_usage(
thread_id=thread_id,
github_login=profile_login,
user_email=user_email,
model_id=model_id,
effort=profile_effort,
source=source,
)
except Exception:
logger.debug("Failed to record agent usage for thread %s", thread_id, exc_info=True)
repo_custom_instructions = await _resolve_repo_custom_instructions(prompt_default_repo)
observability_tools = await _load_observability_tools(
await _observability_authorized(config, profile_login)
)
corridor_tools = await _load_corridor_mcp_tools()
currents_tools: list[Any] = []
notion_tools: list[Any] = []
if profile_login:
try:
currents_tools = await load_currents_tools(profile_login)
except Exception:
logger.warning("Failed to load Currents tools", exc_info=True)
currents_tools = []
try:
notion_tools = await load_notion_tools(profile_login)
except Exception:
logger.warning("Failed to load Notion tools", exc_info=True)
notion_tools = []
logger.info("Returning agent with sandbox for thread %s", thread_id)
main_model = make_model(model_id, **model_kwargs)
subagent_model = make_model(subagent_model_id, **subagent_model_kwargs)
return create_deep_agent(
model=main_model,
system_prompt=construct_system_prompt(
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
working_dir=work_dir,
linear_project_id=linear_project_id,
linear_issue_number=linear_issue_number,
triggering_user_identity=triggering_user_identity,
create_prs=always_create_prs,
default_repo=prompt_default_repo,
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
plan_mode=plan_mode,
plan_url=dashboard_plan_url(thread_id),
repo_custom_instructions=repo_custom_instructions,
thread_url=dashboard_thread_url(thread_id),
corridor_enabled=bool(corridor_tools),
),
tools=[
http_request,
fetch_url,
web_search,
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
enter_plan_mode,
save_plan,
linear_comment,
linear_create_issue,
linear_delete_issue,
linear_get_issue,
linear_get_issue_comments,
linear_list_teams,
linear_update_issue,
open_pull_request,
request_pr_review,
schedule_thread_wakeup,
slack_read_thread_messages,
slack_thread_reply,
*corridor_tools,
*observability_tools,
*currents_tools,
*notion_tools,
],
subagents=[_general_purpose_subagent(subagent_model)],
backend=backend_factory,
middleware=[
SanitizeToolInputsMiddleware(),
ModelCallLimitMiddleware(run_limit=MODEL_CALL_RECURSION_LIMIT, exit_behavior="end"),
ToolErrorMiddleware(),
ToolArtifactMiddleware(),
refresh_github_proxy_before_model,
check_message_queue_before_model,
feat: add optional Slack Assistants API typing status indicator (#1269) * feat: add optional Slack Assistants API typing status indicator Mirrors OpenClaw's pragmatic approach: instead of rebuilding around assistant_thread_started events, just opt into assistants.threads.setStatus to show 'is thinking…' while the agent is working, and clear it when post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so it can be toggled without touching code. * fix(slack): drop redundant clear, add status heartbeat across model calls - Slack auto-clears the typing indicator on bot post; remove the explicit assistants.threads.setStatus("") call from post_slack_thread_reply. - The indicator expires after ~2 minutes; add a before_model middleware that refreshes it on every model tick so it stays visible across long agent runs. Reuses the existing slack_thread.{channel_id,thread_ts} configurable already plumbed for notify_step_limit. - chat:write is sufficient on the bot token (assistant:write is on the way out per Slack docs); no scope or app-config change required. * feat(slack): contextual status text + rotating loading_messages - set_slack_assistant_status now accepts an optional loading_messages list (capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus payload so Slack rotates through them client-side. - The heartbeat middleware derives a contextual status from the last assistant message's tool calls (e.g. "searching the codebase…" after grep, "running commands…" after execute), falling back to the default "is thinking…" when no tool calls or unknown tool name. - Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the contextual status on each refresh. * fix slack assistant status lifecycle --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 10:21:55 -07:00
SlackAssistantStatusMiddleware(),
notify_step_limit_reached,
SandboxCircuitBreakerMiddleware(),
*fallback_middleware,
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
*plan_mode_middleware,
SanitizeThinkingBlocksMiddleware(),
RepairOrphanedToolCallsMiddleware(),
],
).with_config(config)
traced_agent = traced_graph_factory(get_agent, AGENT_TRACING_PROJECT)