mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 17:23:15 +00:00
* fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit7ee3e05724) * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commitf32e492ab4) * feat: reviews block agenda, sticky headers, accurate diff scroll (#1653) Rework the AI-sorted blocks experience on the PR reviews page into a Google-Docs-style outline: the left sidebar is now a clean number+title agenda with scroll-spy highlighting of the active block; each block shows its title + description (sticky) above its diff; and diff rows are pinned to a uniform height so scroll-to lands precisely via the virtualizer's own geometry instead of an estimate-driven correction loop. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit 0b76afdc955e33805c7623d1502a75a9c7c9c1b7) * fix: jump + ResizeObserver settle for review scroll-to (#1655) Replace smooth-scroll plus frame-count correction loops on the PR reviews page with an instant jump that re-asserts its target via a ResizeObserver (the real "layout settled" signal). Block/file navigation and finding/comment centering now land deterministically as off-screen cards mount, files expand, and annotation cards measure, instead of racing a smooth-scroll animation against height reconciliation. Holds bail on user wheel/touch input and after a short ceiling, and a new navigation cancels the previous hold. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> (cherry picked from commit 7530653bba7774d66a54b8bef0d2bbc25f519942) * fix: purge expired thread_wakeup crons (#1656) * fix: purge expired thread_wakeup crons One-shot wakeup crons set an end_time that stops re-firing but the cron row is never deleted, so dead rows accumulate (86 in prod). Add a purge that deletes thread_wakeup crons past their end_time, called opportunistically before scheduling a new wakeup, plus a one-time backfill script. Conservative: matches only kind=thread_wakeup with a past end_time. * chore: retrigger Open SWE review --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit 9e5a1924ef306269322c31342a1831e57831cfee) * fix: add top padding to sticky review block header (#1660) * fix: add top padding to sticky review block header The sticky per-block header on the reviews page had padding below but none above, so the block number badge sat glued against the top edge when pinned. Add matching top padding for breathing room. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: use py-2 shorthand for review block header padding Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit 23bd4a63fc5ba0fe853babf79ed33feb866cc8b2) * fix: use global tokens for sidebar filter popover border (#1661) The filter popover renders via base-ui Menu.Portal into document.body, outside the .agents-ui container where the --ui-* CSS variables are scoped. As a result border-[var(--ui-border)] resolved to an undefined variable and border-color fell back to currentColor, producing a strong near-black border (separators/hover/labels were similarly off). Switch the portaled popup styling to the same global shadcn tokens the theme/settings popover (SidebarUserMenu) already uses (border-border, bg-border, bg-muted, text-muted-foreground). These are defined at :root so they resolve inside portals too, and match the settings popover. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit 63eb9a08209f683016abf01cdcc548bc5905f158) * fix: preserve dashboard redirect after login (#1668) * fix: preserve dashboard redirect after login Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: cover plan login redirect in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit bc7ce59169b5350da7286164afb83a7b037b528d) * Disable React StrictMode (#1654) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit 6575c327a3ac2b107a6e79a04fa61168d779dbf0) * docs(upstream-sync): add cherry-pick runbook Repo-specific runbook for bringing upstream (langchain-ai/open-swe) commits into the fork: triage-sync discovery, the git cp workflow, the triage ledger, themed-branch layout, and conflict/regression handling. --------- Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com>
96 lines
5.1 KiB
Markdown
96 lines
5.1 KiB
Markdown
# Playwright E2E — the full Slack → implement → PR → reply flow
|
|
|
|
This drives the **whole happy path** through two mock UIs:
|
|
|
|
1. A user asks Open SWE to implement something in a **mock Slack** thread.
|
|
2. The **real agent** runs (via `langgraph dev`): it implements the change in a
|
|
**local temp-dir sandbox**, pushes a branch, and opens a PR on a **fake GitHub**.
|
|
3. It posts the PR link back to the **same Slack thread** — visible in the mock UI.
|
|
|
|
## What is faked vs. real
|
|
|
|
Only the **LLM** and the **external SaaS HTTP boundaries** are faked. All agent
|
|
code runs for real.
|
|
|
|
| Piece | Real or fake |
|
|
| ---------------------------------------------------------------- | -------------------------------------------------------------------------- |
|
|
| Slack webhook → `process_slack_mention` → run dispatch | **real** (`agent.webapp`) |
|
|
| `get_agent`, deepagents loop, tools, middleware, prompt | **real** |
|
|
| `open_pull_request`, `slack_thread_reply` tools | **real** |
|
|
| Sandbox | **real** `local` provider, rooted in a throwaway temp dir |
|
|
| Git remote ("GitHub") | **real git**, a local bare repo the agent clones/pushes |
|
|
| The LLM | **fake** — a scripted model (`fake_llm.py`) emitting a fixed tool sequence |
|
|
| `api.github.com` REST (PR create) + dashboard GitHub OAuth login | **fake** (`/fake-gh/...`), state rendered at `/mock/github` |
|
|
| `slack.com/api` (post message, etc.) | **fake** (`/fake-slack/...`), thread rendered at `/mock/slack` |
|
|
| GitHub App token mint, `api.github.com/user` identity | stubbed (offline) |
|
|
|
|
The fake GitHub/Slack stores are the single source of truth the mock UIs render,
|
|
so what Playwright asserts on is exactly what the real agent produced.
|
|
|
|
## Files
|
|
|
|
- `e2e_env.py` — env + constants set before any `agent.*` import (sandbox=local,
|
|
fake API URLs, isolated `GIT_CONFIG_GLOBAL`, bot-token-only mode).
|
|
- `fake_llm.py` — the scripted `BaseChatModel` (the only faked agent piece).
|
|
- `patches.py` — monkeypatches the boundaries (LLM, GitHub/Slack URLs, token mint).
|
|
- `agent_entrypoint.py` — langgraph `agent` graph: applies patches, re-exports the
|
|
real `traced_agent`.
|
|
- `harness.py` — langgraph `http.app`: the real `agent.webapp` plus the fake
|
|
GitHub/Slack APIs, the mock UIs, and the control/compose endpoints.
|
|
- `fakes.py` — in-memory PR/Slack stores + git seeding of the bare remote.
|
|
- `langgraph.e2e.json` — dev-server config pointing at the two entrypoints above.
|
|
- `static/{slack,github}.html` — the mock Slack/GitHub UIs (external SaaS we can't
|
|
run locally). The dashboard is **not** mocked — it's the real `ui/` app.
|
|
- `global-setup.ts` — builds the real `ui/` SPA (once) so the harness can serve it.
|
|
|
|
## The dashboard — the real `ui/` app
|
|
|
|
The dashboard is **not** mocked. The bot's "Open in Web" link
|
|
(`DASHBOARD_BASE_URL/agents/{thread_id}`) loads the **actual built `ui/` React
|
|
app** — served same-origin from the harness so the session cookie and
|
|
`/dashboard/api/*` calls work without CORS. The signed session cookie is real
|
|
(minted via `/control/login`), so per-user authorization is genuine; the only
|
|
extra fake is the OAuth-token store (an external credential).
|
|
|
|
The UI is built by `global-setup.ts` with `VITE_DASHBOARD_API_BASE_URL` pointed at
|
|
the harness. It builds once; set `E2E_FORCE_UI_BUILD=1` to rebuild (e.g. after a
|
|
UI change or port change). Requires Corepack with `pnpm` enabled.
|
|
|
|
## Run
|
|
|
|
```bash
|
|
cd tests/e2e
|
|
npm install
|
|
npx playwright install chromium
|
|
npx playwright test # boots langgraph dev automatically, then runs
|
|
```
|
|
|
|
Watch it in human time:
|
|
|
|
```bash
|
|
SLOW_MO=700 npx playwright test --headed
|
|
```
|
|
|
|
## Artifacts (replay a run)
|
|
|
|
Every test records a **trace** (DOM-snapshot timeline + network + console + source)
|
|
and a **video**; failures also get a screenshot. Locally they land in
|
|
`test-results/<test>/` and are embedded in `playwright-report/`:
|
|
|
|
```bash
|
|
npx playwright show-report # browse runs; each has a Trace tab
|
|
npx playwright show-trace test-results/<test>/trace.zip # open one trace directly
|
|
```
|
|
|
|
In CI the `Playwright E2E` job uploads both `playwright-report/` and
|
|
`test-results/` as the **playwright-report** artifact on the run. Download it,
|
|
then `npx playwright show-report <unzipped-dir>` (or drag a `trace.zip` onto
|
|
<https://trace.playwright.dev>) to replay.
|
|
|
|
Poke at it by hand (from the repo root):
|
|
|
|
```bash
|
|
uv run langgraph dev --config tests/e2e/langgraph.e2e.json --port 2024 \
|
|
--no-browser --allow-blocking --no-reload
|
|
# open http://127.0.0.1:2024/mock/slack and /mock/github
|
|
```
|