mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-10-01 19:03:18 +00:00
8 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c75e76c05e
|
fix(ui): drive Vercel dashboard-API proxy through Nitro routeRules (#76)
The Vercel deploy 404'd at `/` because PR #75's `vercel-build` script (`scripts/build-vercel-output.mjs`) ran `rm -rf .vercel/output` and rebuilt it from `.output/public`. On Vercel CI, Nitro's Vercel preset auto-activates (from the VERCEL env var) and emits the Build Output API layout to `.vercel/output` itself during `vite build`; the script then clobbered that correct output with a static-only config that could not resolve the SPA `_shell.html` fallback or the server functions, so production returned `404: NOT_FOUND` even though the build was READY. Stop fighting Nitro and drive the proxy through it instead: - Remove `scripts/build-vercel-output.mjs` and the `vercel-build` script; restore the plain `vite build` for both local and Vercel builds. - Add an env-driven Nitro `routeRules` proxy in `vite.config.ts`. Nitro's Vercel preset compiles a plain external-URL `proxy` rule into a CDN-level rewrite in `.vercel/output/config.json` at build time, reading LANGGRAPH_BACKEND_URL (the per-project Vercel env var). Proxy, not redirect, so the osw_session cookie stays first-party (same-origin). The build fails loudly if LANGGRAPH_BACKEND_URL is missing on Vercel; the rule is omitted for plain local/off-Vercel builds (which use the E2E_HARNESS mock proxy). - `vercel.json`: `buildCommand` back to `bun run build`; drop the `.output/public` outputDirectory so Vercel serves Nitro's `.vercel/output`. Validated with `LANGGRAPH_BACKEND_URL=... NITRO_PRESET=vercel bun run build`: the generated `.vercel/output/config.json` contains the `/dashboard/api/(.*)` -> backend rewrite ahead of `handle: filesystem` and the SPA catch-all, `_shell.html` and the 699 hashed assets are emitted, and typecheck passes. Supersedes the broken approach in #75. |
||
|
|
3a0e2b4672
|
feat: plan mode with model-driven entry and collaborative review (#1580)
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
|
||
|
|
e4d737e18c
|
fix: optimize agent git diff rendering (#1564)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
be2fe7131b
|
feat: live reviewer eval logs on a dedicated admin page (#1527)
* feat: stream reviewer eval logs on a dedicated admin page Stream the eval subprocess output into a rolling log_tail and persist it during the run (was only captured at exit), so the live output is visible while the eval runs. Move the eval runner off the admin page onto its own /admin/evals page (linked like Review Style Prompts) with a live log viewer. * chore: drop unrelated SSR-register drift from generated route tree * fix(ui): pre-bundle workbox-window to stop dev re-optimize reload The PWA service worker (devOptions.enabled) pulls workbox-window, which Vite discovers after first render and re-optimizes, forcing a reload that cancels in-flight code-split route imports (Failed to fetch dynamically imported module). Pre-bundling it via optimizeDeps.include avoids the mid-session reload. * fix(ui): suppress html hydration warning for pre-hydration theme script The inline theme script sets class="dark"/color-scheme on <html> before React hydrates, so the prerendered HTML never matches. suppressHydrationWarning on <html> silences the (expected) one-level attribute mismatch. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
6f330f777d
|
fix: tidy agents home + mobile logo + PWA dev/icon issues (#1419)
* fix: tidy agents home, mobile logo, and PWA dev/icon issues - Remove recent-runs cards from the new agent page - Scale the ASCII logo to fit narrow viewports instead of overflowing - Serve manifest in dev and skip SW registration in dev to clear console errors - Use a flat, opaque apple-touch-icon so iOS stops adding a black border - Ignore generated ui/dev-dist * feat: add send button and prevent iOS focus-zoom on chat input - Add a circular send button (accent, spinner while sending) to the prompt bar - Bump textarea to 16px on mobile so iOS doesn't zoom the viewport on focus * fix: polish prompt bar and center logo - Center the ASCII logo at all widths - Prevent iOS focus-zoom via viewport maximum-scale instead of bumping the input to 16px, so mobile text stays the intended size - Give the model picker a chip/chevron treatment to anchor it next to the send button * fix: simplify model picker to plain text + chevron * fix: tighten prompt bar padding and shrink send button --------- Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com> |
||
|
|
02cfdbda5b
|
feat: make UI installable as a PWA (#1417)
* feat: make UI installable as a PWA Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com> * fix(pwa): address review feedback - Regenerate bun.lock for new PWA deps (Vercel builds with bun). - Switch SW to prompt registration so deploys don't reload tabs mid-run. - Remove orphaned public/manifest.json; app links the generated /manifest.webmanifest. - Drop json from workbox globPatterns to avoid precaching stray JSON. Verified bun build emits sw.js, manifest.webmanifest, and precaches _shell.html. * fix(pwa): bun.lock, prompt SW registration, tighten globPatterns - Regenerate bun.lock for new PWA deps (Vercel builds with bun). - Switch SW to prompt registration so deploys don't reload tabs mid-run. - Drop json from workbox globPatterns to avoid precaching stray JSON. Verified bun build emits sw.js, manifest.webmanifest, and precaches _shell.html. --------- Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com> |
||
|
|
64e1f75ecc
|
chore(ui): deploy dashboard as a static SPA instead of SSR (#1395)
Enable TanStack Start SPA mode so the build prerenders a static shell (/_shell.html) and emits a fully static bundle under .output/public, removing the Nitro serverless function from the Vercel deploy. The dashboard is a thin client (all data via client-side fetch to the FastAPI /dashboard/api/*), so SSR rendered nothing of value. Point Vercel at the static output and add a SPA catch-all rewrite to the shell for client-side routing, keeping the API proxy rewrite first. |
||
|
|
88856a04fa
|
feat: open-swe dashboard for per-user profile config (#1302)
* feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering: - GitHub App OAuth login → JWT cookie session (cross-domain ready) - profile CRUD against LangGraph Store with model+effort validation - admin gate via CONFIGURED_ADMINS - /repos via /user/installations using the user's encrypted OAuth token CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the Vercel-hosted frontend can call the LangSmith deployment with credentials. * feat: apply dashboard profile model/effort overrides in get_agent Look up the triggering user's GitHub login from config (direct field or GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store, and apply default_model + reasoning_effort to make_model when both are valid. Effort 'max' is captured on the profile but not yet wired through — the OpenAI Reasoning Literal doesn't accept it. * feat: ui/ TanStack Start dashboard for profile config Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template, base-ui primitives, Tailwind v4). Three routes: - /login — Sign in with GitHub (links to /dashboard/api/auth/login) - /profile — Edit default model, reasoning effort, default repo - /admin — Admin-only: list users and edit other profiles API client (src/lib/api.ts) uses credentials: include so the osw_session cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL points at the LangSmith deployment. Effort options re-render when the model changes; 'max' on Opus 4.7 is captured on the profile but ignored downstream until anthropic reasoning is wired through make_model. * feat: searchable Combobox for default repo picker Replaces the Select with a base-ui Combobox so users can filter by typing, the popup is wider than the trigger so full owner/repo names are readable, and the list caps at max-h-80 to stay on screen. * fix: address review comments + wire default_repo and Anthropic thinking Security/correctness fixes from PR review: * Open redirect: validate `redirect_to` in `/auth/login` against `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it into the state JWT. Anything off-allowlist falls back to the dashboard base URL. (PR #1302 r3250054386) * Login CSRF: bind the OAuth `state` to the requesting browser. At `/auth/login` we generate a fresh nonce, set it as a short-lived HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback` we require the cookie nonce to hash-match the state JWT's nonce_hash (constant-time compare). (PR #1302 r3250054395) * RMW race in profile vs token writes: split storage into two namespaces — `["profiles"]` for user-editable settings and `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now only writes its own namespace so an in-flight profile save can no longer clobber a fresh token from a concurrent re-login (and vice versa). (PR #1302 r3250054393) * /repos pagination: follow `Link: rel="next"` for both `/user/installations` and per-installation `/repositories` with per_page=100, capped at 1000 items. (PR #1302 r3250054401) Feature wires: * default_repo: applied as a fallback in `get_slack_repo_config` (after explicit-repo / thread metadata, before the env defaults) and in the Linear webhook (after comment-body extraction, before team mapping). Both paths resolve the triggering user's GitHub login via GITHUB_USER_EMAIL_MAP and read the profile's default_repo. * Anthropic "thinking" effort: `make_model` now accepts a `thinking` kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max} to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is anthropic. OpenAI path still ignores "max" since the Literal doesn't accept it. |