mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 13:53:15 +00:00
6 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f87847baa4
|
feat: port plan-review & workflow-approval UX (#159)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: port plan-review & workflow-approval UX (#135)
Port six upstream commits onto dev:
-
|
||
|
|
a53a96d37b
|
feat: editable plan mode — owner hand-edits the plan before approval (#1610, #80) (#130)
* feat: editable plan mode — owner hand-edits the plan before approval (#1610, #80) Adds a PUT /dashboard/api/plan/{thread_id} endpoint and PlanReview UI edit mode so the thread owner can refine the published plan markdown by hand. The edited markdown is re-published as "ready" (preserving reviewer comments), mirrored into the sandbox plan.md, and handed to the agent as the source of truth on approve. Approve now reads the published plan content strictly (raise_on_error=True) so a transient store failure aborts instead of silently dropping an edited plan, matching the comment-read contract. The banner-overlap fix (collapsed git-panel clearing the "Review plan" link) was already ported in #128; this picks up the remaining edit-mode pieces. Refs #80 * fix(plan): make approve_plan idempotent, dispatch before persisting, fix comment count SH-128-03: approve_plan set status APPROVED before dispatching the follow-up run and had no already-approved guard, so a failed dispatch left the plan stuck approved-but-undispatched and a double-submit double-dispatched + double-posted the Slack notice. And the Slack notice counted len(comments) including empty comments _format_comments filters out. - Return 409 when the plan is already approved (idempotent double-click/retry). - Dispatch the implementation run BEFORE persisting APPROVED so a dispatch failure leaves the plan re-approvable. _dispatch_followup passes plan_mode explicitly, so the run is unaffected by the reorder. - Count only non-empty comments in the Slack approval notice. Fixed here (not on #128) because #128's approve_plan is rewritten on this branch; #129 inherits it. Adds tests for the 409, the filtered count, and the dispatch-before-status ordering. --------- Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com> |
||
|
|
b3b0274403
|
feat: Re-land deferred upstream features on modular webhooks (#80) (#128)
* fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect). |
||
|
|
2b01652754
|
refactor: adopt modular webhook architecture (#1621) + port fork customizations (#85)
* Adopt upstream modular webhook skeleton (#1621) Apply the durable-interrupt-dispatch refactor: split the monolithic webapp.py into a thin routing layer plus per-source handlers in webhooks/{github,slack,linear}.py, and add completion.py, dispatch.py, and reconcile.py. Reconcile fork divergence by keeping the Bedrock/ Fireworks cross-provider fallback, the no-agent-attribution prompt policy, the dashboard-handoff re-export, and the Slack channel-info cache. ci_autofix is restored on the new dispatch model in a later commit. Refs: #80 * Port fork webhook security delta onto modular handlers Re-apply the fork's security customizations that #1621 did not carry: Linear webhook replay protection (freshness window on the signed webhookTimestamp), per-repo token-cache binding threaded through the thread token resolvers, the INTERNAL_BOT_LOGINS self-check in the review-finding-reply path, and a user-mapping cache refresh before email resolution on the issue and PR-comment paths (multi-replica staleness). Existing fork security tests pass unchanged. Refs: #80 * Restore CI auto-fix on the modular dispatch model Bring back ci_autofix.py and the ci_monitor graph that #1621 deleted, re-wiring the fork's security-reviewed PR-babysitting onto the new structure: the CI-event, autofix-toggle, and review-feedback handlers move into webhooks/github.py and the github_webhook router re-gains the check_run/check_suite/workflow_run/status routing plus the autofix command and actionable-review branches. Auto-fix runs now dispatch through dispatch_agent_run (durability + completion webhook) while keeping the deliberate batch-while-busy skip-rule via get_thread_active_status. Restore langgraph.json's ci_monitor entry and the fork autofix tests (dispatch mock + import paths re-pointed). Refs: #80 * Reformat and update docs for the modular webhook split Point CLAUDE.md and deploy/MIGRATION.md at the new webhooks/ modules and the dispatch/completion/reconcile contract, and mark the user-mapping cache-refresh fix as applied on the GitHub handlers. Refs: #80 * Restore reject backstop for autofix dispatch A burst of near-simultaneous CI events for one head SHA can slip past the busy-check before the dedupe SHA is recorded, so dispatch the autofix path with multitask_strategy=reject (dev's prior platform default) to drop duplicate concurrent creates instead of letting them interrupt each other. Also make the completion failure-reply dedup claim-then-post and drop the unreachable interrupted branch. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
ca9280d25c
|
refactor(open-swe): plain HTTP comments instead of Yjs/BlockNote collab (#1601)
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
* refactor(plan-mode): replace Yjs/BlockNote collab with plain HTTP comments
Drop the realtime collaborative editor (it can't work behind Vercel's
rewrite — WebSocket upgrades aren't proxied to the external LangGraph
backend) in favor of a simple whole-document comments API over plain HTTP.
Backend:
- Remove the Yjs WebSocket server (plan_collab.py), its lifespan, and the
collab router; drop pycrdt / pycrdt-websocket deps.
- plan_store: replace the Yjs snapshot with comment CRUD (one store item per
comment under ["plan","comments",thread_id]).
- plan_api: add GET/POST/DELETE comment endpoints; approve/reject now read
comments server-side and format them for the follow-up run (no longer
client-harvested). Comment delete is author-or-owner; approve stays owner-only.
Frontend:
- PlanReview renders the plan markdown read-only and shows a comments panel
(list + add, polled every 4s for cross-user visibility).
- Drop @blocknote/*, y-websocket, yjs; lib/plan exposes get/add/deletePlanComment.
Tests: unit tests for the comments API + route registration; e2e drives the
HTTP comment UI (owner + collaborator, cross-user visibility, owner-only approve,
PR echoes the harvested feedback).
* fix(open-swe): clear stale plan comments on republish; fail loud on store errors
Address reviewer feedback:
- Clear comments when a revised plan is published (save_plan_content) so
feedback on the prior revision doesn't resurface and get re-fed to the agent.
- list_plan_comments gains raise_on_error; approve/reject read comments before
mutating state and propagate store failures (500) instead of silently
dispatching the follow-up run with no feedback.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
|
||
|
|
3a0e2b4672
|
feat: plan mode with model-driven entry and collaborative review (#1580)
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
|