mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 08:03:15 +00:00
* chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
976 lines
38 KiB
Python
976 lines
38 KiB
Python
"""Main entry point and CLI loop for Open SWE agent."""
|
|
# ruff: noqa: E402
|
|
|
|
# Suppress deprecation warnings from langchain_core (e.g., Pydantic V1 on Python 3.14+)
|
|
# ruff: noqa: E402
|
|
import logging
|
|
import os
|
|
import time
|
|
import warnings
|
|
from collections.abc import Sequence
|
|
from typing import Any
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
from langgraph.graph.state import RunnableConfig
|
|
from langgraph.pregel import Pregel
|
|
from langgraph_sdk import get_client
|
|
|
|
warnings.filterwarnings("ignore", module="langchain_core._api.deprecation")
|
|
|
|
import asyncio
|
|
|
|
# Suppress Pydantic v1 compatibility warnings from langchain on Python 3.14+
|
|
warnings.filterwarnings("ignore", message=".*Pydantic V1.*", category=UserWarning)
|
|
|
|
from deepagents import create_deep_agent
|
|
from deepagents.backends import LangSmithSandbox
|
|
from deepagents.backends.protocol import SandboxBackendProtocol
|
|
from deepagents.middleware.subagents import GENERAL_PURPOSE_SUBAGENT, SubAgent
|
|
from langchain.agents.middleware import ModelCallLimitMiddleware
|
|
from langchain_core.language_models import BaseChatModel
|
|
from langsmith.sandbox import SandboxClientError
|
|
|
|
from .dashboard.admin import is_observability_authorized
|
|
from .dashboard.agent_overrides import (
|
|
load_profile,
|
|
normalize_profile_overrides,
|
|
normalize_profile_subagent_overrides,
|
|
profile_author_prs_as_user,
|
|
profile_create_prs,
|
|
resolve_github_login,
|
|
)
|
|
from .dashboard.agent_usage import record_agent_thread_usage
|
|
from .dashboard.options import DEFAULT_MODEL_ID, SUPPORTED_MODEL_IDS, model_supports_effort
|
|
from .dashboard.repo_snapshots import resolve_repo_snapshot_id
|
|
from .dashboard.team_settings import (
|
|
get_team_default_model_pair,
|
|
get_team_default_repo,
|
|
)
|
|
from .dashboard.user_mappings import email_for_login
|
|
from .integrations.corridor_mcp import load_corridor_tools
|
|
from .integrations.currents_tools import load_currents_tools
|
|
from .integrations.datadog_mcp import load_datadog_tools
|
|
from .integrations.langsmith import _configure_github_proxy
|
|
from .integrations.langsmith_tools import load_langsmith_tools
|
|
from .integrations.notion_mcp import load_notion_tools
|
|
from .middleware import (
|
|
ModelFallbackMiddleware,
|
|
PlanModeMiddleware,
|
|
SandboxCircuitBreakerMiddleware,
|
|
SanitizeThinkingBlocksMiddleware,
|
|
SanitizeToolInputsMiddleware,
|
|
SlackAssistantStatusMiddleware,
|
|
ToolArtifactMiddleware,
|
|
ToolErrorMiddleware,
|
|
WorkflowPushGuardMiddleware,
|
|
check_message_queue_before_model,
|
|
ensure_no_empty_msg,
|
|
notify_step_limit_reached,
|
|
refresh_github_proxy_before_model,
|
|
)
|
|
from .prompt import construct_system_prompt
|
|
from .tools import (
|
|
enter_plan_mode,
|
|
fetch_url,
|
|
http_request,
|
|
linear_comment,
|
|
linear_create_issue,
|
|
linear_delete_issue,
|
|
linear_get_issue,
|
|
linear_get_issue_comments,
|
|
linear_list_teams,
|
|
linear_update_issue,
|
|
open_pull_request,
|
|
request_pr_review,
|
|
save_plan,
|
|
schedule_thread_wakeup,
|
|
slack_read_thread_messages,
|
|
slack_thread_reply,
|
|
web_search,
|
|
)
|
|
from .utils.auth import resolve_github_token
|
|
from .utils.authorship import (
|
|
OPEN_SWE_BOT_EMAIL,
|
|
OPEN_SWE_BOT_NAME,
|
|
resolve_triggering_user_identity,
|
|
)
|
|
from .utils.dashboard_links import dashboard_plan_url, dashboard_thread_url
|
|
from .utils.github_app import (
|
|
BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS,
|
|
RUNTIME_PROXY_TOKEN_PERMISSIONS,
|
|
PermissionMap,
|
|
get_github_app_installation_token_with_expiry,
|
|
)
|
|
from .utils.github_proxy import record_proxy_token_expiry
|
|
from .utils.github_token import repo_cache_key
|
|
from .utils.model import (
|
|
fallback_model_id_for,
|
|
make_model,
|
|
provider_model_kwargs,
|
|
)
|
|
from .utils.sandbox import create_sandbox
|
|
from .utils.sandbox_paths import aresolve_sandbox_work_dir
|
|
from .utils.tracing import AGENT_TRACING_PROJECT, traced_graph_factory
|
|
|
|
client = get_client()
|
|
|
|
SANDBOX_CREATING = "__creating__"
|
|
SANDBOX_CREATION_TIMEOUT = 180
|
|
SANDBOX_POLL_INTERVAL = 1.0
|
|
|
|
from .utils.sandbox_state import (
|
|
SANDBOX_BACKENDS,
|
|
get_bound_repo_from_metadata,
|
|
get_sandbox_id_from_metadata,
|
|
set_sandbox_backend,
|
|
unwrap_sandbox_backend,
|
|
)
|
|
|
|
|
|
async def _resolve_prompt_default_repo(configurable: dict[str, Any]) -> dict[str, str] | None:
|
|
repo_config = configurable.get("repo")
|
|
if isinstance(repo_config, dict):
|
|
owner = repo_config.get("owner")
|
|
name = repo_config.get("name")
|
|
if isinstance(owner, str) and isinstance(name, str):
|
|
return {"owner": owner, "name": name}
|
|
|
|
if configurable.get("repo_explicitly_none") is True:
|
|
return None
|
|
|
|
try:
|
|
return await get_team_default_repo()
|
|
except Exception:
|
|
logger.debug("Failed to load team default repo for prompt", exc_info=True)
|
|
return None
|
|
|
|
|
|
async def _resolve_repo_custom_instructions(
|
|
default_repo: dict[str, str] | None,
|
|
) -> str | None:
|
|
"""Load per-repo custom agent instructions for the resolved default repo."""
|
|
if not default_repo or not default_repo.get("owner") or not default_repo.get("name"):
|
|
return None
|
|
try:
|
|
from .dashboard.agent_instructions import get_repo_agent_instructions
|
|
|
|
return await get_repo_agent_instructions(default_repo["owner"], default_repo["name"])
|
|
except Exception:
|
|
logger.debug("Failed to load repo custom agent instructions", exc_info=True)
|
|
return None
|
|
|
|
|
|
async def _start_langsmith_sandbox_if_needed(sandbox_backend: SandboxBackendProtocol) -> None:
|
|
"""Start a LangSmith sandbox before operations that require it to be running."""
|
|
if os.getenv("SANDBOX_TYPE", "langsmith") != "langsmith":
|
|
return
|
|
current_backend = unwrap_sandbox_backend(sandbox_backend)
|
|
if not isinstance(current_backend, LangSmithSandbox):
|
|
return
|
|
|
|
sandbox = current_backend._sandbox # noqa: SLF001
|
|
status = await asyncio.to_thread(sandbox._client.get_sandbox_status, sandbox.name) # noqa: SLF001
|
|
status_name = getattr(status, "status", status)
|
|
status_name = getattr(status_name, "value", status_name)
|
|
status_text = str(status_name or "").lower()
|
|
if status_text in {"running", "ready"}:
|
|
return
|
|
|
|
logger.info(
|
|
"Starting LangSmith sandbox %s before proxy refresh (status=%s)",
|
|
current_backend.id,
|
|
status_text or "unknown",
|
|
)
|
|
await asyncio.to_thread(sandbox.start)
|
|
|
|
|
|
async def _resolve_proxy_token(
|
|
github_proxy_token: str | None,
|
|
*,
|
|
permissions: PermissionMap | None = None,
|
|
) -> tuple[str | None, str | None, PermissionMap | None]:
|
|
"""Resolve the proxy token, its expiry, and the effective permission scope."""
|
|
if github_proxy_token:
|
|
return github_proxy_token, None, None
|
|
if permissions is not None:
|
|
token, expires_at = await get_github_app_installation_token_with_expiry(
|
|
permissions=permissions
|
|
)
|
|
return token, expires_at, permissions
|
|
|
|
token, expires_at = await get_github_app_installation_token_with_expiry(
|
|
permissions=RUNTIME_PROXY_TOKEN_PERMISSIONS,
|
|
log_errors=False,
|
|
)
|
|
if token:
|
|
return token, expires_at, RUNTIME_PROXY_TOKEN_PERMISSIONS
|
|
|
|
logger.warning("Retrying GitHub proxy token mint without optional Actions read permission")
|
|
token, expires_at = await get_github_app_installation_token_with_expiry(
|
|
permissions=BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS
|
|
)
|
|
return token, expires_at, BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS if token else None
|
|
|
|
|
|
async def _resolve_snapshot_id_for_repo(repo: dict[str, str] | None) -> str | None:
|
|
"""Resolve a repo's ready snapshot id; ``None`` falls back to the default.
|
|
|
|
Never raises: any failure resolves to ``None`` so sandbox creation falls
|
|
back to the configured ``DEFAULT_SANDBOX_SNAPSHOT_ID``.
|
|
"""
|
|
if not repo:
|
|
return None
|
|
try:
|
|
return await resolve_repo_snapshot_id(repo.get("owner"), repo.get("name"))
|
|
except Exception: # noqa: BLE001
|
|
logger.debug("Failed to resolve repo-scoped snapshot", exc_info=True)
|
|
return None
|
|
|
|
|
|
async def _create_sandbox_with_proxy(
|
|
github_proxy_token: str | None = None,
|
|
*,
|
|
thread_id: str | None = None,
|
|
github_proxy_repositories: Sequence[str] | None = None,
|
|
repo: dict[str, str] | None = None,
|
|
) -> SandboxBackendProtocol:
|
|
"""Create a new sandbox with GitHub proxy auth configured."""
|
|
snapshot_id = await _resolve_snapshot_id_for_repo(repo)
|
|
sandbox_backend = await asyncio.to_thread(create_sandbox, snapshot_id=snapshot_id)
|
|
|
|
sandbox_type = os.getenv("SANDBOX_TYPE", "langsmith")
|
|
if sandbox_type == "langsmith":
|
|
token, expires_at, permissions = await _resolve_proxy_token(github_proxy_token)
|
|
if not token:
|
|
msg = "Cannot configure proxy: GitHub App installation token is unavailable"
|
|
logger.error(msg)
|
|
raise ValueError(msg)
|
|
await _start_langsmith_sandbox_if_needed(sandbox_backend)
|
|
await asyncio.to_thread(_configure_github_proxy, sandbox_backend.id, token)
|
|
record_proxy_token_expiry(
|
|
thread_id,
|
|
expires_at,
|
|
repositories=github_proxy_repositories,
|
|
permissions=permissions,
|
|
)
|
|
|
|
return sandbox_backend
|
|
|
|
|
|
async def _refresh_github_proxy(
|
|
sandbox_backend: SandboxBackendProtocol,
|
|
github_proxy_token: str | None = None,
|
|
*,
|
|
thread_id: str | None = None,
|
|
github_proxy_repositories: Sequence[str] | None = None,
|
|
) -> None:
|
|
"""Refresh GitHub proxy credentials for reused LangSmith sandboxes."""
|
|
if os.getenv("SANDBOX_TYPE", "langsmith") != "langsmith":
|
|
return
|
|
|
|
token, expires_at, permissions = await _resolve_proxy_token(github_proxy_token)
|
|
if not token:
|
|
logger.warning(
|
|
"Skipping GitHub proxy refresh for sandbox %s: installation token unavailable",
|
|
sandbox_backend.id,
|
|
)
|
|
return
|
|
|
|
current_backend = unwrap_sandbox_backend(sandbox_backend)
|
|
await _start_langsmith_sandbox_if_needed(current_backend)
|
|
await asyncio.to_thread(_configure_github_proxy, current_backend.id, token)
|
|
record_proxy_token_expiry(
|
|
thread_id,
|
|
expires_at,
|
|
repositories=github_proxy_repositories,
|
|
permissions=permissions,
|
|
)
|
|
|
|
|
|
async def _refresh_github_proxy_or_recreate(
|
|
sandbox_backend: SandboxBackendProtocol,
|
|
thread_id: str,
|
|
github_proxy_token: str | None = None,
|
|
github_proxy_repositories: Sequence[str] | None = None,
|
|
repo: dict[str, str] | None = None,
|
|
) -> SandboxBackendProtocol:
|
|
"""Refresh proxy credentials, recreating stale LangSmith sandboxes on failure."""
|
|
try:
|
|
await _refresh_github_proxy(
|
|
sandbox_backend,
|
|
github_proxy_token,
|
|
thread_id=thread_id,
|
|
github_proxy_repositories=github_proxy_repositories,
|
|
)
|
|
except Exception: # noqa: BLE001
|
|
logger.warning(
|
|
"Failed to refresh GitHub proxy for sandbox %s on thread %s, recreating sandbox",
|
|
sandbox_backend.id,
|
|
thread_id,
|
|
exc_info=True,
|
|
)
|
|
return await _recreate_sandbox(
|
|
thread_id,
|
|
github_proxy_token=github_proxy_token,
|
|
github_proxy_repositories=github_proxy_repositories,
|
|
repo=repo,
|
|
)
|
|
return sandbox_backend
|
|
|
|
|
|
async def _configure_git_identity(sandbox_backend: SandboxBackendProtocol) -> None:
|
|
await asyncio.to_thread(
|
|
sandbox_backend.execute,
|
|
f"git config --global user.name '{OPEN_SWE_BOT_NAME}' && "
|
|
f"git config --global user.email '{OPEN_SWE_BOT_EMAIL}'",
|
|
)
|
|
|
|
|
|
async def _recreate_sandbox(
|
|
thread_id: str,
|
|
*,
|
|
github_proxy_token: str | None = None,
|
|
github_proxy_repositories: Sequence[str] | None = None,
|
|
repo: dict[str, str] | None = None,
|
|
) -> SandboxBackendProtocol:
|
|
"""Recreate a sandbox after a connection failure.
|
|
|
|
Sets the SANDBOX_CREATING sentinel and creates a fresh sandbox
|
|
(with proxy auth configured), swapping the per-thread proxy target.
|
|
The agent is responsible for cloning repos via tools.
|
|
"""
|
|
await client.threads.update(thread_id=thread_id, metadata=_creating_metadata())
|
|
try:
|
|
sandbox_backend = set_sandbox_backend(
|
|
thread_id,
|
|
await _create_sandbox_with_proxy(
|
|
github_proxy_token,
|
|
thread_id=thread_id,
|
|
github_proxy_repositories=github_proxy_repositories,
|
|
repo=repo,
|
|
),
|
|
)
|
|
except Exception:
|
|
logger.exception("Failed to recreate sandbox after connection failure")
|
|
await client.threads.update(thread_id=thread_id, metadata=_RESET_METADATA)
|
|
raise
|
|
return sandbox_backend
|
|
|
|
|
|
async def check_or_recreate_sandbox(
|
|
sandbox_backend: SandboxBackendProtocol,
|
|
thread_id: str,
|
|
github_proxy_token: str | None = None,
|
|
github_proxy_repositories: Sequence[str] | None = None,
|
|
repo: dict[str, str] | None = None,
|
|
) -> SandboxBackendProtocol:
|
|
"""Check if a cached sandbox is reachable; recreate it if not.
|
|
|
|
Pings the sandbox with a lightweight command. If the sandbox is
|
|
unreachable (SandboxClientError), it is torn down and a fresh one
|
|
is created via _recreate_sandbox.
|
|
|
|
Returns the original backend if healthy, or a new one if recreated.
|
|
"""
|
|
try:
|
|
await asyncio.to_thread(sandbox_backend.execute, "echo ok")
|
|
except SandboxClientError:
|
|
logger.warning(
|
|
"Cached sandbox is no longer reachable for thread %s, recreating",
|
|
thread_id,
|
|
)
|
|
sandbox_backend = await _recreate_sandbox(
|
|
thread_id,
|
|
github_proxy_token=github_proxy_token,
|
|
github_proxy_repositories=github_proxy_repositories,
|
|
repo=repo,
|
|
)
|
|
return sandbox_backend
|
|
|
|
|
|
def _creating_metadata() -> dict[str, Any]:
|
|
"""Metadata that claims the cross-process creation lock with a timestamp."""
|
|
return {"sandbox_id": SANDBOX_CREATING, "sandbox_creating_at": time.time()}
|
|
|
|
|
|
_RESET_METADATA: dict[str, Any] = {"sandbox_id": None, "sandbox_creating_at": None}
|
|
|
|
|
|
async def _resolve_creating_sentinel(thread_id: str) -> str | None:
|
|
"""Resolve a ``__creating__`` sentinel seen with no cached backend.
|
|
|
|
The sentinel is a cross-process lock: another worker may still be creating
|
|
the sandbox. Poll live thread metadata until it resolves to a real id. Only
|
|
when the sentinel is older than ``SANDBOX_CREATION_TIMEOUT`` (e.g. the
|
|
creating worker was restarted) is it treated as stale: metadata is reset and
|
|
``None`` is returned so the caller creates a fresh sandbox. A sentinel with
|
|
no timestamp (written before this field existed) is also treated as stale.
|
|
"""
|
|
while True:
|
|
thread = await client.threads.get(thread_id)
|
|
metadata = thread.get("metadata", {}) if isinstance(thread, dict) else {}
|
|
sandbox_id = metadata.get("sandbox_id") if isinstance(metadata, dict) else None
|
|
|
|
if sandbox_id != SANDBOX_CREATING:
|
|
return sandbox_id if isinstance(sandbox_id, str) else None
|
|
|
|
creating_at = metadata.get("sandbox_creating_at") if isinstance(metadata, dict) else None
|
|
age = time.time() - creating_at if isinstance(creating_at, (int, float)) else None
|
|
if age is None or age > SANDBOX_CREATION_TIMEOUT:
|
|
logger.warning(
|
|
"Resetting stale SANDBOX_CREATING for thread %s (age=%s)", thread_id, age
|
|
)
|
|
await client.threads.update(thread_id=thread_id, metadata=_RESET_METADATA)
|
|
return None
|
|
|
|
await asyncio.sleep(SANDBOX_POLL_INTERVAL)
|
|
|
|
|
|
def graph_loaded_for_execution(config: RunnableConfig) -> bool:
|
|
"""Check if the graph is loaded for actual execution vs introspection."""
|
|
return (
|
|
config["configurable"].get("__is_for_execution__", False)
|
|
if "configurable" in config
|
|
else False
|
|
)
|
|
|
|
|
|
class SandboxRepoMismatchError(RuntimeError):
|
|
"""Raised when a thread_id is presented for a repo it is not bound to.
|
|
|
|
A thread is bound to exactly one repo. A different repo presenting a
|
|
colliding thread_id (e.g. an attacker-named branch whose first UUID matches
|
|
another thread) must never reuse this thread's sandbox or token.
|
|
"""
|
|
|
|
def __init__(self, thread_id: str, bound_repo: str, current_repo: str) -> None:
|
|
self.thread_id = thread_id
|
|
self.bound_repo = bound_repo
|
|
self.current_repo = current_repo
|
|
super().__init__(
|
|
f"Thread {thread_id} is bound to repo {bound_repo}, "
|
|
f"refusing to serve sandbox for {current_repo}"
|
|
)
|
|
|
|
|
|
async def ensure_sandbox_for_thread(
|
|
thread_id: str,
|
|
*,
|
|
github_proxy_token: str | None = None,
|
|
github_proxy_repositories: Sequence[str] | None = None,
|
|
repo: dict[str, str] | None = None,
|
|
) -> SandboxBackendProtocol:
|
|
"""Get-or-create a healthy sandbox bound to ``thread_id``.
|
|
|
|
Implements the four-state lifecycle described in AGENTS.md:
|
|
|
|
1. Cached in memory → ping; recreate on ``SandboxClientError``.
|
|
2. Metadata says ``__creating__`` and no cache → wait for the creating
|
|
worker; only reset if the sentinel is proven stale (timestamp/timeout).
|
|
3. No sandbox at all → create one and persist the id.
|
|
4. Metadata has an id but no cache → reconnect; recreate on failure.
|
|
|
|
For LangSmith sandboxes, also refreshes the GitHub App proxy auth. When
|
|
``repo`` has a ``ready`` repo-scoped snapshot, newly created sandboxes boot
|
|
from it; otherwise the configured ``DEFAULT_SANDBOX_SNAPSHOT_ID`` is used.
|
|
Persists the resulting ``sandbox_id`` to thread metadata, and on the
|
|
first creation/reconnect for this thread initializes git identity.
|
|
"""
|
|
sandbox_backend = SANDBOX_BACKENDS.get(thread_id)
|
|
sandbox_id = await get_sandbox_id_from_metadata(thread_id)
|
|
|
|
if sandbox_id == SANDBOX_CREATING and not sandbox_backend:
|
|
logger.info("Sandbox creation in progress for thread %s, waiting...", thread_id)
|
|
sandbox_id = await _resolve_creating_sentinel(thread_id)
|
|
|
|
# Repo-binding guard (TID-COLLIDE-01): a sandbox is never served to a repo
|
|
# unless its binding is known and matches.
|
|
current_repo = repo_cache_key(repo)
|
|
bound_repo = await get_bound_repo_from_metadata(thread_id)
|
|
proxy_bound = getattr(sandbox_backend, "bound_repo", None)
|
|
effective_bound = bound_repo or (proxy_bound if isinstance(proxy_bound, str) else None)
|
|
if current_repo and effective_bound and effective_bound != current_repo:
|
|
# Known binding that does not match the current repo: refuse outright so a
|
|
# colliding thread_id from a different repo cannot reuse/clobber it.
|
|
logger.error(
|
|
"Repo mismatch for thread %s: bound=%s current=%s; refusing sandbox reuse",
|
|
thread_id,
|
|
effective_bound,
|
|
current_repo,
|
|
)
|
|
raise SandboxRepoMismatchError(thread_id, effective_bound, current_repo)
|
|
if (
|
|
current_repo
|
|
and not effective_bound
|
|
and sandbox_backend is None
|
|
and isinstance(sandbox_id, str)
|
|
and sandbox_id not in (None, SANDBOX_CREATING)
|
|
):
|
|
# Fail CLOSED for unbound-legacy threads (migration window): a thread with a
|
|
# persisted sandbox_id but no in-memory cache and no recorded bound_repo
|
|
# cannot be confirmed to belong to the current repo, so never
|
|
# reconnect-and-serve it. Drop the stale id and recreate a fresh sandbox
|
|
# bound to this repo below.
|
|
logger.error(
|
|
"reconnect-with-missing-binding for thread %s: persisted sandbox %s has no "
|
|
"bound_repo; refusing reuse and recreating for repo %s",
|
|
thread_id,
|
|
sandbox_id,
|
|
current_repo,
|
|
)
|
|
sandbox_id = None
|
|
|
|
if sandbox_backend:
|
|
logger.info("Using cached sandbox backend for thread %s", thread_id)
|
|
original_sandbox_id = sandbox_backend.id
|
|
sandbox_backend = await check_or_recreate_sandbox(
|
|
sandbox_backend, thread_id, github_proxy_token, github_proxy_repositories, repo
|
|
)
|
|
if sandbox_backend.id == original_sandbox_id:
|
|
sandbox_backend = await _refresh_github_proxy_or_recreate(
|
|
sandbox_backend, thread_id, github_proxy_token, github_proxy_repositories, repo
|
|
)
|
|
elif sandbox_id is None:
|
|
logger.info("Creating new sandbox for thread %s", thread_id)
|
|
await client.threads.update(thread_id=thread_id, metadata=_creating_metadata())
|
|
try:
|
|
sandbox_backend = await _create_sandbox_with_proxy(
|
|
github_proxy_token,
|
|
thread_id=thread_id,
|
|
github_proxy_repositories=github_proxy_repositories,
|
|
repo=repo,
|
|
)
|
|
logger.info("Sandbox created: %s", sandbox_backend.id)
|
|
except Exception:
|
|
logger.exception("Failed to create sandbox")
|
|
try:
|
|
await client.threads.update(thread_id=thread_id, metadata=_RESET_METADATA)
|
|
except Exception:
|
|
logger.exception("Failed to reset sandbox_id metadata")
|
|
raise
|
|
else:
|
|
logger.info("Connecting to existing sandbox %s", sandbox_id)
|
|
created_replacement_sandbox = False
|
|
try:
|
|
sandbox_backend = await asyncio.to_thread(create_sandbox, sandbox_id)
|
|
except Exception:
|
|
logger.warning("Failed to connect to existing sandbox %s, creating new one", sandbox_id)
|
|
await client.threads.update(thread_id=thread_id, metadata=_creating_metadata())
|
|
try:
|
|
sandbox_backend = await _create_sandbox_with_proxy(
|
|
github_proxy_token,
|
|
thread_id=thread_id,
|
|
github_proxy_repositories=github_proxy_repositories,
|
|
repo=repo,
|
|
)
|
|
created_replacement_sandbox = True
|
|
except Exception:
|
|
logger.exception("Failed to create replacement sandbox")
|
|
await client.threads.update(thread_id=thread_id, metadata=_RESET_METADATA)
|
|
raise
|
|
if not created_replacement_sandbox:
|
|
original_sandbox_id = sandbox_backend.id
|
|
sandbox_backend = await check_or_recreate_sandbox(
|
|
sandbox_backend, thread_id, github_proxy_token, github_proxy_repositories, repo
|
|
)
|
|
if sandbox_backend.id == original_sandbox_id:
|
|
sandbox_backend = await _refresh_github_proxy_or_recreate(
|
|
sandbox_backend, thread_id, github_proxy_token, github_proxy_repositories, repo
|
|
)
|
|
|
|
sandbox_backend = set_sandbox_backend(thread_id, sandbox_backend, repo=current_repo)
|
|
|
|
metadata_update: dict[str, Any] = {}
|
|
if sandbox_id != sandbox_backend.id:
|
|
metadata_update["sandbox_id"] = sandbox_backend.id
|
|
if current_repo and bound_repo != current_repo:
|
|
metadata_update["bound_repo"] = current_repo
|
|
if metadata_update:
|
|
await client.threads.update(thread_id=thread_id, metadata=metadata_update)
|
|
|
|
# Re-apply git identity every run: cached/reconnected sandboxes may have
|
|
# lost their `--global` config (or had it overwritten), and Vercel preview
|
|
# deploys reject commits whose author email can't be resolved to a GitHub
|
|
# account.
|
|
await _configure_git_identity(sandbox_backend)
|
|
|
|
return sandbox_backend
|
|
|
|
|
|
DEFAULT_LLM_MODEL_ID = DEFAULT_MODEL_ID
|
|
DEFAULT_LLM_MAX_TOKENS = 64_000
|
|
DEFAULT_RECURSION_LIMIT = 9_999
|
|
# High cap to support long-running tasks; a run that hits it still ends with a
|
|
# signal via notify_step_limit_reached rather than dying silently.
|
|
MODEL_CALL_RECURSION_LIMIT = 5_000
|
|
|
|
# Mutating external tools hidden from the model while plan mode is active so it
|
|
# can only research and propose a plan. File edit tools stay available so the
|
|
# agent can draft and revise a plan under `/workspace/plans/`; prompt guidance
|
|
# restricts them to that plan file outside cloned repositories. `execute` stays available;
|
|
# plan-mode shell discipline (no mutating commands) is instructed via the system
|
|
# prompt rather than enforced. `http_request` is excluded because it can
|
|
# POST/PUT/PATCH/DELETE to external services — read-only web research goes
|
|
# through `web_search` / `fetch_url`. `task` is excluded because the
|
|
# general-purpose subagent is built with its own filesystem/PR/Linear tools and
|
|
# does not inherit this exclusion, so delegating to it would bypass the read-only
|
|
# intent.
|
|
PLAN_MODE_EXCLUDED_TOOLS: frozenset[str] = frozenset(
|
|
{
|
|
"write_file",
|
|
"edit_file",
|
|
"task",
|
|
"http_request",
|
|
"open_pull_request",
|
|
"request_pr_review",
|
|
"linear_create_issue",
|
|
"linear_update_issue",
|
|
"linear_delete_issue",
|
|
}
|
|
)
|
|
|
|
|
|
def _general_purpose_subagent(model: BaseChatModel) -> SubAgent:
|
|
return {
|
|
"name": GENERAL_PURPOSE_SUBAGENT["name"],
|
|
"description": GENERAL_PURPOSE_SUBAGENT["description"],
|
|
"system_prompt": GENERAL_PURPOSE_SUBAGENT["system_prompt"],
|
|
"model": model,
|
|
}
|
|
|
|
|
|
def _get_cached_sandbox_backend(thread_id: str) -> SandboxBackendProtocol:
|
|
sandbox_backend = SANDBOX_BACKENDS.get(thread_id)
|
|
if sandbox_backend is None:
|
|
raise RuntimeError(f"No sandbox backend cached for thread {thread_id}")
|
|
return sandbox_backend
|
|
|
|
|
|
async def _observability_authorized(config: RunnableConfig, profile_login: str | None) -> bool:
|
|
"""Whether the triggering user may use the team observability tools.
|
|
|
|
Gates on admin / explicitly-authorized emails so prompt-injected runs from
|
|
untrusted contributors cannot reach the team's Datadog/LangSmith data.
|
|
"""
|
|
configurable = (config or {}).get("configurable") or {}
|
|
slack_thread = configurable.get("slack_thread") or {}
|
|
config_login = configurable.get("github_login")
|
|
candidate_login = profile_login or (config_login if isinstance(config_login, str) else None)
|
|
candidate_emails = [
|
|
configurable.get("user_email"),
|
|
slack_thread.get("triggering_user_email"),
|
|
]
|
|
if any(is_observability_authorized(email, login=candidate_login) for email in candidate_emails):
|
|
return True
|
|
return is_observability_authorized(
|
|
await email_for_login(candidate_login), login=candidate_login
|
|
)
|
|
|
|
|
|
async def _load_observability_tools(authorized: bool) -> list[Any]:
|
|
"""Datadog (MCP) + LangSmith read tools when the team has connected them.
|
|
|
|
Credentials live server-side in team settings; the sandbox never holds them.
|
|
Only loaded for authorized (admin / allow-listed) triggering users so an
|
|
untrusted run cannot exfiltrate team observability data. Failures degrade to
|
|
no tools so the agent still starts.
|
|
"""
|
|
if not authorized:
|
|
return []
|
|
try:
|
|
datadog_tools, langsmith_tools = await asyncio.gather(
|
|
load_datadog_tools(),
|
|
load_langsmith_tools(),
|
|
)
|
|
except Exception:
|
|
logger.warning("Failed to load observability tools", exc_info=True)
|
|
return []
|
|
return [*datadog_tools, *langsmith_tools]
|
|
|
|
|
|
async def _load_corridor_mcp_tools() -> list[Any]:
|
|
"""Corridor MCP tools when the deployment environment has configured them."""
|
|
try:
|
|
return await load_corridor_tools()
|
|
except Exception:
|
|
logger.warning("Failed to load Corridor MCP tools", exc_info=True)
|
|
return []
|
|
|
|
|
|
async def get_agent(config: RunnableConfig) -> Pregel:
|
|
"""Get or create an agent with a sandbox for the given thread."""
|
|
thread_id = config["configurable"].get("thread_id", None)
|
|
|
|
config["recursion_limit"] = DEFAULT_RECURSION_LIMIT
|
|
|
|
if thread_id is None or not graph_loaded_for_execution(config):
|
|
logger.info("No thread_id or not for execution, returning agent without sandbox")
|
|
return create_deep_agent(
|
|
system_prompt="",
|
|
tools=[],
|
|
).with_config(config)
|
|
|
|
github_token, _expires_at = await resolve_github_token(config, thread_id)
|
|
profile_login = resolve_github_login(config)
|
|
configurable = (config or {}).get("configurable") or {}
|
|
prompt_default_repo = await _resolve_prompt_default_repo(configurable)
|
|
|
|
# Commit identity must follow the SAME default-bot decision as the token
|
|
# (SH-IDSPLIT-01): by default slack/dashboard/schedule runs author commits as the
|
|
# app bot, so resolve the triggering USER's git identity ONLY when authoring as the
|
|
# user (the author_prs_as_user opt-in, or a non-default source). Otherwise leave it
|
|
# None so construct_system_prompt sets the bot identity (OPEN_SWE_BOT_NAME/EMAIL) and
|
|
# commits don't get mis-attributed to a human who didn't write them.
|
|
if configurable.get("source") in ("slack", "dashboard", "schedule"):
|
|
_author_as_user = bool(
|
|
isinstance(profile_login, str)
|
|
and profile_login.strip()
|
|
and profile_author_prs_as_user(await load_profile(profile_login.strip()))
|
|
)
|
|
else:
|
|
_author_as_user = True
|
|
|
|
async def _no_triggering_identity() -> Any:
|
|
return None
|
|
|
|
if _author_as_user:
|
|
triggering_user_identity_task = asyncio.create_task(
|
|
asyncio.to_thread(resolve_triggering_user_identity, config, github_token)
|
|
)
|
|
else:
|
|
triggering_user_identity_task = asyncio.create_task(_no_triggering_identity())
|
|
sandbox_task = asyncio.create_task(
|
|
ensure_sandbox_for_thread(thread_id, repo=prompt_default_repo)
|
|
)
|
|
team_defaults_task = asyncio.create_task(get_team_default_model_pair("agent"))
|
|
profile_task = asyncio.create_task(load_profile(profile_login)) if profile_login else None
|
|
try:
|
|
triggering_user_identity, sandbox_backend, team_defaults = await asyncio.gather(
|
|
triggering_user_identity_task,
|
|
sandbox_task,
|
|
team_defaults_task,
|
|
)
|
|
except SandboxRepoMismatchError as exc:
|
|
# Repo-binding refusal at the run boundary: log for alarming and surface the
|
|
# already-sanitized terminal error (no sandbox/token internals) to the caller,
|
|
# rather than letting an opaque deep-stack exception crash-loop the worker.
|
|
logger.error("Refusing agent run for thread %s: %s", thread_id, exc)
|
|
for pending in (triggering_user_identity_task, team_defaults_task, profile_task):
|
|
if pending is not None and not pending.done():
|
|
pending.cancel()
|
|
raise RuntimeError(str(exc)) from exc
|
|
profile = await profile_task if profile_task is not None else None
|
|
del github_token
|
|
|
|
linear_issue = config["configurable"].get("linear_issue", {})
|
|
linear_project_id = linear_issue.get("linear_project_id", "")
|
|
linear_issue_number = linear_issue.get("linear_issue_number", "")
|
|
|
|
work_dir = await aresolve_sandbox_work_dir(sandbox_backend)
|
|
|
|
def backend_factory(_runtime: object, _thread_id: str = thread_id) -> SandboxBackendProtocol:
|
|
return _get_cached_sandbox_backend(_thread_id)
|
|
|
|
(model_id, profile_effort), (subagent_model_id, subagent_effort) = team_defaults
|
|
logger.info("Using team default agent model: model=%s effort=%s", model_id, profile_effort)
|
|
logger.info(
|
|
"Using team default agent subagent model: model=%s effort=%s",
|
|
subagent_model_id,
|
|
subagent_effort,
|
|
)
|
|
|
|
if profile_login and profile:
|
|
overridden_model, overridden_effort = normalize_profile_overrides(profile)
|
|
if overridden_model:
|
|
logger.info(
|
|
"Applying dashboard profile override for %s: model=%s effort=%s",
|
|
profile_login,
|
|
overridden_model,
|
|
overridden_effort,
|
|
)
|
|
model_id = overridden_model
|
|
profile_effort = overridden_effort
|
|
subagent_model_id = overridden_model
|
|
subagent_effort = overridden_effort
|
|
overridden_subagent_model, overridden_subagent_effort = (
|
|
normalize_profile_subagent_overrides(profile)
|
|
)
|
|
if overridden_subagent_model:
|
|
logger.info(
|
|
"Applying dashboard profile subagent override for %s: model=%s effort=%s",
|
|
profile_login,
|
|
overridden_subagent_model,
|
|
overridden_subagent_effort,
|
|
)
|
|
subagent_model_id = overridden_subagent_model
|
|
subagent_effort = overridden_subagent_effort
|
|
|
|
per_thread_model = configurable.get("agent_model_id")
|
|
per_thread_effort = configurable.get("agent_effort")
|
|
if (
|
|
isinstance(per_thread_model, str)
|
|
and per_thread_model in SUPPORTED_MODEL_IDS
|
|
and isinstance(per_thread_effort, str)
|
|
and model_supports_effort(per_thread_model, per_thread_effort)
|
|
):
|
|
logger.info(
|
|
"Applying per-thread model override: model=%s effort=%s",
|
|
per_thread_model,
|
|
per_thread_effort,
|
|
)
|
|
model_id = per_thread_model
|
|
profile_effort = per_thread_effort
|
|
subagent_model_id = per_thread_model
|
|
subagent_effort = per_thread_effort
|
|
|
|
always_create_prs = profile_create_prs(profile)
|
|
if always_create_prs:
|
|
logger.info("Always Create PRs enabled by profile for %s", profile_login)
|
|
|
|
model_kwargs = provider_model_kwargs(
|
|
model_id,
|
|
profile_effort,
|
|
max_tokens=DEFAULT_LLM_MAX_TOKENS,
|
|
)
|
|
subagent_model_kwargs = provider_model_kwargs(
|
|
subagent_model_id,
|
|
subagent_effort,
|
|
max_tokens=DEFAULT_LLM_MAX_TOKENS,
|
|
)
|
|
|
|
fallback_model_id = os.environ.get("LLM_FALLBACK_MODEL_ID") or fallback_model_id_for(model_id)
|
|
fallback_middleware: list[Any] = []
|
|
if fallback_model_id and fallback_model_id != model_id:
|
|
fallback_kwargs = provider_model_kwargs(
|
|
fallback_model_id, None, max_tokens=DEFAULT_LLM_MAX_TOKENS
|
|
)
|
|
fallback_middleware.append(
|
|
ModelFallbackMiddleware(make_model(fallback_model_id, **fallback_kwargs))
|
|
)
|
|
logger.info("Configured model fallback %s -> %s", model_id, fallback_model_id)
|
|
|
|
# Plan mode is entered only when the model decides to (the `enter_plan_mode`
|
|
# tool sets it in run state). The configurable value just carries that
|
|
# decision across a thread's messages and the approve/reject follow-ups; a
|
|
# fresh run with nothing set starts out of plan mode.
|
|
plan_mode = configurable.get("plan_mode") is True
|
|
if plan_mode:
|
|
logger.info("Plan mode enabled for thread %s", thread_id)
|
|
# Installed unconditionally and state-aware: it also restricts tools after a
|
|
# mid-run `enter_plan_mode` call, not just when plan mode is set up front.
|
|
plan_mode_middleware: list[Any] = [
|
|
PlanModeMiddleware(excluded=PLAN_MODE_EXCLUDED_TOOLS, initial=plan_mode)
|
|
]
|
|
|
|
source = (
|
|
configurable.get("source") if isinstance(configurable.get("source"), str) else "dashboard"
|
|
)
|
|
user_email = configurable.get("user_email")
|
|
user_email = user_email if isinstance(user_email, str) else ""
|
|
try:
|
|
await client.threads.update(
|
|
thread_id=thread_id,
|
|
metadata={
|
|
"agent_kind": "agent",
|
|
"model": model_id,
|
|
"effort": profile_effort,
|
|
"source": source,
|
|
"plan_mode": plan_mode,
|
|
},
|
|
)
|
|
await record_agent_thread_usage(
|
|
thread_id=thread_id,
|
|
github_login=profile_login,
|
|
user_email=user_email,
|
|
model_id=model_id,
|
|
effort=profile_effort,
|
|
source=source,
|
|
)
|
|
except Exception:
|
|
logger.debug("Failed to record agent usage for thread %s", thread_id, exc_info=True)
|
|
|
|
repo_custom_instructions = await _resolve_repo_custom_instructions(prompt_default_repo)
|
|
|
|
observability_tools = await _load_observability_tools(
|
|
await _observability_authorized(config, profile_login)
|
|
)
|
|
corridor_tools = await _load_corridor_mcp_tools()
|
|
|
|
currents_tools: list[Any] = []
|
|
notion_tools: list[Any] = []
|
|
if profile_login:
|
|
try:
|
|
currents_tools = await load_currents_tools(profile_login)
|
|
except Exception:
|
|
logger.warning("Failed to load Currents tools", exc_info=True)
|
|
currents_tools = []
|
|
try:
|
|
notion_tools = await load_notion_tools(profile_login)
|
|
except Exception:
|
|
logger.warning("Failed to load Notion tools", exc_info=True)
|
|
notion_tools = []
|
|
|
|
logger.info("Returning agent with sandbox for thread %s", thread_id)
|
|
main_model = make_model(model_id, **model_kwargs)
|
|
subagent_model = make_model(subagent_model_id, **subagent_model_kwargs)
|
|
return create_deep_agent(
|
|
model=main_model,
|
|
system_prompt=construct_system_prompt(
|
|
working_dir=work_dir,
|
|
linear_project_id=linear_project_id,
|
|
linear_issue_number=linear_issue_number,
|
|
triggering_user_identity=triggering_user_identity,
|
|
create_prs=always_create_prs,
|
|
default_repo=prompt_default_repo,
|
|
plan_mode=plan_mode,
|
|
plan_url=dashboard_plan_url(thread_id),
|
|
repo_custom_instructions=repo_custom_instructions,
|
|
thread_url=dashboard_thread_url(thread_id),
|
|
corridor_enabled=bool(corridor_tools),
|
|
),
|
|
tools=[
|
|
http_request,
|
|
fetch_url,
|
|
web_search,
|
|
enter_plan_mode,
|
|
save_plan,
|
|
linear_comment,
|
|
linear_create_issue,
|
|
linear_delete_issue,
|
|
linear_get_issue,
|
|
linear_get_issue_comments,
|
|
linear_list_teams,
|
|
linear_update_issue,
|
|
open_pull_request,
|
|
request_pr_review,
|
|
schedule_thread_wakeup,
|
|
slack_read_thread_messages,
|
|
slack_thread_reply,
|
|
*corridor_tools,
|
|
*observability_tools,
|
|
*currents_tools,
|
|
*notion_tools,
|
|
],
|
|
subagents=[_general_purpose_subagent(subagent_model)],
|
|
backend=backend_factory,
|
|
middleware=[
|
|
SanitizeToolInputsMiddleware(),
|
|
ModelCallLimitMiddleware(run_limit=MODEL_CALL_RECURSION_LIMIT, exit_behavior="end"),
|
|
ToolErrorMiddleware(),
|
|
ToolArtifactMiddleware(),
|
|
WorkflowPushGuardMiddleware(),
|
|
refresh_github_proxy_before_model,
|
|
check_message_queue_before_model,
|
|
SlackAssistantStatusMiddleware(),
|
|
ensure_no_empty_msg,
|
|
notify_step_limit_reached,
|
|
SandboxCircuitBreakerMiddleware(),
|
|
*fallback_middleware,
|
|
*plan_mode_middleware,
|
|
SanitizeThinkingBlocksMiddleware(),
|
|
],
|
|
).with_config(config)
|
|
|
|
|
|
traced_agent = traced_graph_factory(get_agent, AGENT_TRACING_PROJECT)
|