open-swe/agent/dashboard/plan_api.py
Ramon Nogueira ca9280d25c
refactor(open-swe): plain HTTP comments instead of Yjs/BlockNote collab (#1601)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

* refactor(plan-mode): replace Yjs/BlockNote collab with plain HTTP comments

Drop the realtime collaborative editor (it can't work behind Vercel's
rewrite — WebSocket upgrades aren't proxied to the external LangGraph
backend) in favor of a simple whole-document comments API over plain HTTP.

Backend:
- Remove the Yjs WebSocket server (plan_collab.py), its lifespan, and the
  collab router; drop pycrdt / pycrdt-websocket deps.
- plan_store: replace the Yjs snapshot with comment CRUD (one store item per
  comment under ["plan","comments",thread_id]).
- plan_api: add GET/POST/DELETE comment endpoints; approve/reject now read
  comments server-side and format them for the follow-up run (no longer
  client-harvested). Comment delete is author-or-owner; approve stays owner-only.

Frontend:
- PlanReview renders the plan markdown read-only and shows a comments panel
  (list + add, polled every 4s for cross-user visibility).
- Drop @blocknote/*, y-websocket, yjs; lib/plan exposes get/add/deletePlanComment.

Tests: unit tests for the comments API + route registration; e2e drives the
HTTP comment UI (owner + collaborator, cross-user visibility, owner-only approve,
PR echoes the harvested feedback).

* fix(open-swe): clear stale plan comments on republish; fail loud on store errors

Address reviewer feedback:
- Clear comments when a revised plan is published (save_plan_content) so
  feedback on the prior revision doesn't resurface and get re-fed to the agent.
- list_plan_comments gains raise_on_error; approve/reject read comments before
  mutating state and propagate store failures (500) instead of silently
  dispatching the follow-up run with no feedback.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 22:12:42 +00:00

222 lines
8.3 KiB
Python

"""REST API for the plan-review page: read the plan, comment, approve, or request
changes — all plain HTTP, no CRDT/WebSocket.
Reviewers leave whole-document comments via this API; they're stored server-side
and listed for everyone who can read the thread. On approve/reject the comments
are read back here, formatted, and handed to the agent as the instruction for the
follow-up run. The agent never sees comments during review — only this aggregated
feedback at the decision point.
Permissions: any authenticated org member can read a surfaced thread, comment, and
request changes (reject); only the thread owner can approve. A comment can be
deleted by its author or the thread owner.
"""
from __future__ import annotations
import logging
from typing import Any
from fastapi import APIRouter, Depends, HTTPException
from langgraph_sdk import get_client
from pydantic import BaseModel
from .oauth import require_same_origin_for_mutations, require_session
from .plan_store import (
PLAN_STATUS_APPROVED,
PLAN_STATUS_REVISING,
add_plan_comment,
delete_plan_comment,
get_plan_content,
list_plan_comments,
set_plan_status,
)
from .thread_api import (
_repo_config_from_metadata,
_thread_is_readable,
_thread_source,
_user_owns_thread,
)
logger = logging.getLogger(__name__)
plan_router = APIRouter(
prefix="/dashboard/api/plan",
tags=["plan"],
dependencies=[Depends(require_same_origin_for_mutations)],
)
_SESSION_DEP = Depends(require_session)
class CommentBody(BaseModel):
body: str
async def _thread_metadata(thread_id: str) -> dict[str, Any]:
client = get_client()
try:
thread = await client.threads.get(thread_id)
except Exception as exc: # noqa: BLE001
raise HTTPException(404, "thread not found") from exc
metadata = (
thread.get("metadata") if isinstance(thread, dict) else getattr(thread, "metadata", None)
)
return metadata if isinstance(metadata, dict) else {}
@plan_router.get("/{thread_id}")
async def get_plan(thread_id: str, session: dict[str, Any] = _SESSION_DEP) -> dict[str, Any]:
metadata = await _thread_metadata(thread_id)
if not _thread_is_readable(metadata):
raise HTTPException(404, "thread not found")
login = session["sub"]
email = session.get("email")
content = await get_plan_content(thread_id) or {}
return {
"threadId": thread_id,
"status": content.get("status") or metadata.get("plan_status") or "planning",
"markdown": content.get("markdown", ""),
"isOwner": _user_owns_thread(metadata, login, email),
"user": {
"id": login,
"login": login,
"email": email,
"name": session.get("name") or login,
},
}
@plan_router.get("/{thread_id}/comments")
async def get_plan_comments(
thread_id: str, session: dict[str, Any] = _SESSION_DEP
) -> dict[str, Any]:
metadata = await _thread_metadata(thread_id)
if not _thread_is_readable(metadata):
raise HTTPException(404, "thread not found")
return {"comments": await list_plan_comments(thread_id)}
@plan_router.post("/{thread_id}/comments")
async def post_plan_comment(
thread_id: str, body: CommentBody, session: dict[str, Any] = _SESSION_DEP
) -> dict[str, Any]:
metadata = await _thread_metadata(thread_id)
if not _thread_is_readable(metadata):
raise HTTPException(404, "thread not found")
text = body.body.strip()
if not text:
raise HTTPException(422, "comment body cannot be empty")
login = session["sub"]
return await add_plan_comment(
thread_id, author=session.get("name") or login, author_login=login, body=text
)
@plan_router.delete("/{thread_id}/comments/{comment_id}")
async def remove_plan_comment(
thread_id: str, comment_id: str, session: dict[str, Any] = _SESSION_DEP
) -> dict[str, Any]:
metadata = await _thread_metadata(thread_id)
if not _thread_is_readable(metadata):
raise HTTPException(404, "thread not found")
comments = await list_plan_comments(thread_id)
target = next((c for c in comments if c.get("id") == comment_id), None)
if target is None:
raise HTTPException(404, "comment not found")
login = session["sub"]
is_owner = _user_owns_thread(metadata, login, session.get("email"))
if target.get("author_login") != login and not is_owner:
raise HTTPException(403, "only the author or the plan owner can delete a comment")
await delete_plan_comment(thread_id, comment_id)
return {"ok": True}
@plan_router.post("/{thread_id}/approve")
async def approve_plan(thread_id: str, session: dict[str, Any] = _SESSION_DEP) -> dict[str, Any]:
metadata = await _thread_metadata(thread_id)
if not _user_owns_thread(metadata, session["sub"], session.get("email")):
raise HTTPException(403, "only the plan owner can approve")
# Read comments BEFORE mutating state: a store failure here aborts the
# decision (500) rather than dispatching the run without the feedback.
feedback = _format_comments(await list_plan_comments(thread_id, raise_on_error=True))
await set_plan_status(thread_id, PLAN_STATUS_APPROVED, plan_mode=False)
if feedback:
text = (
"The plan has been approved. Implement it now, taking this reviewer "
f"feedback into account:\n\n{feedback}"
)
else:
text = "The plan has been approved. Implement it now as described in the plan."
await _dispatch_followup(thread_id, metadata, text, plan_mode=False)
return {"status": PLAN_STATUS_APPROVED}
@plan_router.post("/{thread_id}/reject")
async def reject_plan(thread_id: str, session: dict[str, Any] = _SESSION_DEP) -> dict[str, Any]:
metadata = await _thread_metadata(thread_id)
if not _thread_is_readable(metadata):
raise HTTPException(404, "thread not found")
feedback = _format_comments(await list_plan_comments(thread_id, raise_on_error=True))
await set_plan_status(thread_id, PLAN_STATUS_REVISING, plan_mode=True)
text = (
"The plan needs changes before implementation. Address this reviewer "
"feedback and publish an updated plan with the save_plan tool:\n\n"
f"{feedback or '(no specific comments were left)'}"
)
await _dispatch_followup(thread_id, metadata, text, plan_mode=True)
return {"status": PLAN_STATUS_REVISING}
def _format_comments(comments: list[dict[str, Any]]) -> str:
lines: list[str] = []
index = 1
for comment in comments:
body = str(comment.get("body", "")).strip()
if not body:
continue
author = str(comment.get("author") or "reviewer").strip()
lines.append(f"{index}. {author}: {body}")
index += 1
return "\n".join(lines)
async def _dispatch_followup(
thread_id: str, metadata: dict[str, Any], text: str, *, plan_mode: bool
) -> None:
"""Continue the existing thread with a new instruction run.
Runs on the same LangGraph thread, so the agent resumes from the checkpoint
with the full planning history plus this instruction. The configurable is
rebuilt from the thread's stored owner/repo/Slack context so the agent can
push, open a PR, and reply in the original channel.
"""
configurable: dict[str, Any] = {
"thread_id": thread_id,
"source": _thread_source(metadata) or "slack",
}
email = metadata.get("triggering_user_email")
if isinstance(email, str) and email:
configurable["user_email"] = email
login = metadata.get("github_login")
if isinstance(login, str) and login:
configurable["github_login"] = login
repo = _repo_config_from_metadata(metadata)
if repo:
configurable["repo"] = repo
source_context = metadata.get("source_context")
if isinstance(source_context, dict):
slack_thread = source_context.get("slack_thread")
if isinstance(slack_thread, dict):
configurable["slack_thread"] = slack_thread
# Carry the decision to the follow-up run: approve continues out of plan
# mode (implement), reject stays in plan mode (revise the plan).
configurable["plan_mode"] = plan_mode
client = get_client()
await client.runs.create(
thread_id,
"agent",
input={"messages": [{"role": "user", "content": text}]},
config={"configurable": configurable},
if_not_exists="create",
)