open-swe/agent/dashboard/plan_api.py
Ramon Nogueira 3a0e2b4672
feat: plan mode with model-driven entry and collaborative review (#1580)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:06:58 -07:00

188 lines
6.6 KiB
Python

"""REST API for the plan-review page: read the plan, approve, or request changes.
Comment threads themselves live in the collaborative Yjs document (BlockNote's
``YjsThreadStore``) and are harvested client-side; on approve/reject the client
sends the harvested feedback here, where it is formatted and handed to the agent
as the instruction for the follow-up run. The agent never sees comments during
review — only this aggregated feedback at the decision point.
Permissions: any authenticated org member can read a surfaced thread and request
changes (reject); only the thread owner can approve.
"""
from __future__ import annotations
import logging
from typing import Any
from fastapi import APIRouter, Depends, HTTPException
from langgraph_sdk import get_client
from pydantic import BaseModel, Field
from .oauth import require_same_origin_for_mutations, require_session
from .plan_store import (
PLAN_STATUS_APPROVED,
PLAN_STATUS_REVISING,
get_plan_content,
set_plan_status,
)
from .thread_api import (
_repo_config_from_metadata,
_thread_is_readable,
_thread_source,
_user_owns_thread,
)
logger = logging.getLogger(__name__)
plan_router = APIRouter(
prefix="/dashboard/api/plan",
tags=["plan"],
dependencies=[Depends(require_same_origin_for_mutations)],
)
_SESSION_DEP = Depends(require_session)
class PlanComment(BaseModel):
author: str | None = None
body: str = ""
quote: str | None = None
resolved: bool = False
class PlanDecisionBody(BaseModel):
comments: list[PlanComment] = Field(default_factory=list)
async def _thread_metadata(thread_id: str) -> dict[str, Any]:
client = get_client()
try:
thread = await client.threads.get(thread_id)
except Exception as exc: # noqa: BLE001
raise HTTPException(404, "thread not found") from exc
metadata = (
thread.get("metadata") if isinstance(thread, dict) else getattr(thread, "metadata", None)
)
return metadata if isinstance(metadata, dict) else {}
@plan_router.get("/{thread_id}")
async def get_plan(thread_id: str, session: dict[str, Any] = _SESSION_DEP) -> dict[str, Any]:
metadata = await _thread_metadata(thread_id)
if not _thread_is_readable(metadata):
raise HTTPException(404, "thread not found")
login = session["sub"]
email = session.get("email")
content = await get_plan_content(thread_id) or {}
return {
"threadId": thread_id,
"status": content.get("status") or metadata.get("plan_status") or "planning",
"markdown": content.get("markdown", ""),
"isOwner": _user_owns_thread(metadata, login, email),
"user": {
"id": login,
"login": login,
"email": email,
"name": session.get("name") or login,
},
}
@plan_router.post("/{thread_id}/approve")
async def approve_plan(
thread_id: str,
body: PlanDecisionBody,
session: dict[str, Any] = _SESSION_DEP,
) -> dict[str, Any]:
metadata = await _thread_metadata(thread_id)
if not _user_owns_thread(metadata, session["sub"], session.get("email")):
raise HTTPException(403, "only the plan owner can approve")
await set_plan_status(thread_id, PLAN_STATUS_APPROVED, plan_mode=False)
feedback = _format_comments(body.comments)
if feedback:
text = (
"The plan has been approved. Implement it now, taking this reviewer "
f"feedback into account:\n\n{feedback}"
)
else:
text = "The plan has been approved. Implement it now as described in the plan."
await _dispatch_followup(thread_id, metadata, text, plan_mode=False)
return {"status": PLAN_STATUS_APPROVED}
@plan_router.post("/{thread_id}/reject")
async def reject_plan(
thread_id: str,
body: PlanDecisionBody,
session: dict[str, Any] = _SESSION_DEP,
) -> dict[str, Any]:
metadata = await _thread_metadata(thread_id)
if not _thread_is_readable(metadata):
raise HTTPException(404, "thread not found")
await set_plan_status(thread_id, PLAN_STATUS_REVISING, plan_mode=True)
feedback = _format_comments(body.comments)
text = (
"The plan needs changes before implementation. Address this reviewer "
"feedback and publish an updated plan with the save_plan tool:\n\n"
f"{feedback or '(no specific comments were left)'}"
)
await _dispatch_followup(thread_id, metadata, text, plan_mode=True)
return {"status": PLAN_STATUS_REVISING}
def _format_comments(comments: list[PlanComment]) -> str:
lines: list[str] = []
for index, comment in enumerate(comments, start=1):
body = comment.body.strip()
if not body:
continue
author = (comment.author or "reviewer").strip()
quote = (comment.quote or "").strip()
status = " (resolved)" if comment.resolved else ""
if quote:
lines.append(f'{index}. On "{quote}"{status}:\n - {author}: {body}')
else:
lines.append(f"{index}. {author}{status}: {body}")
return "\n".join(lines)
async def _dispatch_followup(
thread_id: str, metadata: dict[str, Any], text: str, *, plan_mode: bool
) -> None:
"""Continue the existing thread with a new instruction run.
Runs on the same LangGraph thread, so the agent resumes from the checkpoint
with the full planning history plus this instruction. The configurable is
rebuilt from the thread's stored owner/repo/Slack context so the agent can
push, open a PR, and reply in the original channel.
"""
configurable: dict[str, Any] = {
"thread_id": thread_id,
"source": _thread_source(metadata) or "slack",
}
email = metadata.get("triggering_user_email")
if isinstance(email, str) and email:
configurable["user_email"] = email
login = metadata.get("github_login")
if isinstance(login, str) and login:
configurable["github_login"] = login
repo = _repo_config_from_metadata(metadata)
if repo:
configurable["repo"] = repo
source_context = metadata.get("source_context")
if isinstance(source_context, dict):
slack_thread = source_context.get("slack_thread")
if isinstance(slack_thread, dict):
configurable["slack_thread"] = slack_thread
# Carry the decision to the follow-up run: approve continues out of plan
# mode (implement), reject stays in plan mode (revise the plan).
configurable["plan_mode"] = plan_mode
client = get_client()
await client.runs.create(
thread_id,
"agent",
input={"messages": [{"role": "user", "content": text}]},
config={"configurable": configurable},
if_not_exists="create",
)