open-swe/agent/tools/save_plan.py

117 lines
4.4 KiB
Python
Raw Normal View History

feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
"""Tool: ``save_plan``. Publish the sandbox plan file for review.
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
Reads the Markdown plan file the agent created in the sandbox and publishes it to
the plan-review page, where the user and collaborators read it, comment inline,
and approve or request changes. Available in plan mode (it does not modify the
repository under review).
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
"""
from __future__ import annotations
import logging
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
from collections.abc import Mapping
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
from typing import Any
from langgraph.config import get_config
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
from ..dashboard.plan_store import PLAN_FILE_DIRECTORY, PLAN_STATUS_READY, save_plan_content
from ..utils.sandbox_state import get_sandbox_backend
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
logger = logging.getLogger(__name__)
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
_MAX_PLAN_LINES = 20_000
_MARKDOWN_EXTENSIONS = (".md", ".markdown")
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
async def save_plan(plan_file_path: str) -> dict[str, Any]:
"""Publish a Markdown plan file from the sandbox for review.
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
Use this in plan mode once your plan is ready. First create a Markdown file
under ``/workspace/plans/`` using a dated, descriptive filename, then pass
that file path here. The file contents are published to the plan-review page
linked in the conversation, where the user (the owner) and any collaborators
can read it, leave inline comments, and then approve it or request changes.
Call it again to publish a revised file when addressing feedback.
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
Write the plan in standard Markdown — headings, bullet/numbered lists, and
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
fenced code blocks all render. Keep it concise and high level, focusing on
approach, decisions/tradeoffs, risks, and verification; avoid file/function
details unless they are unusually tricky or controversial.
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
Args:
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
plan_file_path: Path to the Markdown plan file in the sandbox.
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
Returns:
``{success: True, path}`` on success, or ``{success: False, error}``.
"""
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
if not isinstance(plan_file_path, str):
return {"success": False, "error": "plan_file_path must be a string"}
path = plan_file_path.strip()
if not path:
return {"success": False, "error": "plan_file_path cannot be empty"}
if not _is_markdown_path(path):
return {
"success": False,
"error": f"plan_file_path must point to a Markdown file in {PLAN_FILE_DIRECTORY}",
}
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
try:
config = get_config()
except Exception:
config = {}
configurable = config.get("configurable", {}) if isinstance(config, dict) else {}
thread_id = configurable.get("thread_id") if isinstance(configurable, dict) else None
if not thread_id:
return {"success": False, "error": "no thread_id in run config"}
try:
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
content = (await _read_plan_file(str(thread_id), path)).strip()
if not content:
return {"success": False, "error": "plan file cannot be empty"}
await _save(str(thread_id), content, path)
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
except Exception as exc: # noqa: BLE001
logger.exception("save_plan failed for thread %s", thread_id)
return {"success": False, "error": f"failed to save plan: {exc}"}
return {"success": True, "path": path}
feat: Re-land deferred upstream features on modular webhooks (#80) (#128) * fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
async def _save(thread_id: str, content: str, path: str) -> None:
await save_plan_content(
thread_id, markdown=content, status=PLAN_STATUS_READY, plan_file_path=path
)
async def _read_plan_file(thread_id: str, path: str) -> str:
backend = await get_sandbox_backend(thread_id)
result = await backend.aread(path, offset=0, limit=_MAX_PLAN_LINES)
error = _value(result, "error")
if error:
raise ValueError(error)
file_data = _value(result, "file_data")
if file_data is None:
raise ValueError("plan file could not be read")
encoding = _value(file_data, "encoding")
if encoding is not None and encoding != "utf-8":
raise ValueError("plan file must be UTF-8 text")
content = _value(file_data, "content")
if not isinstance(content, str):
raise ValueError("plan file content was not text")
if content.count("\n") + 1 >= _MAX_PLAN_LINES:
raise ValueError("plan file is too large")
return content
def _value(value: Any, key: str) -> Any:
if isinstance(value, Mapping):
return value.get(key)
return getattr(value, key, None)
def _is_markdown_path(path: str) -> bool:
if "\x00" in path or not path.startswith(f"{PLAN_FILE_DIRECTORY}/"):
return False
filename = path.removeprefix(f"{PLAN_FILE_DIRECTORY}/")
if not filename or "/" in filename:
return False
return filename.lower().endswith(_MARKDOWN_EXTENSIONS)