mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 19:43:15 +00:00
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
44 lines
1.7 KiB
TypeScript
44 lines
1.7 KiB
TypeScript
import { defineConfig, devices } from "@playwright/test";
|
|
import { resolve } from "node:path";
|
|
|
|
const repoRoot = resolve(__dirname, "..", "..");
|
|
const PORT = Number(process.env.E2E_PORT ?? 2024);
|
|
const baseURL = `http://127.0.0.1:${PORT}`;
|
|
|
|
export default defineConfig({
|
|
testDir: "./tests",
|
|
globalSetup: "./global-setup.ts",
|
|
fullyParallel: false,
|
|
workers: 1,
|
|
forbidOnly: !!process.env.CI,
|
|
retries: process.env.CI ? 1 : 0,
|
|
timeout: 90_000,
|
|
expect: { timeout: 60_000 },
|
|
reporter: [["list"], ["html", { open: "never" }]],
|
|
use: {
|
|
baseURL,
|
|
// Always capture the replayable artifacts: a trace (DOM snapshots, network,
|
|
// console, source — open with `npx playwright show-trace`) and a screen
|
|
// recording, plus a screenshot on failure. The CI job uploads them.
|
|
trace: "on",
|
|
video: "on",
|
|
screenshot: "only-on-failure",
|
|
// The built UI ships a PWA service worker; block it so tests never hit a
|
|
// stale cache and always see live API responses.
|
|
serviceWorkers: "block",
|
|
// SLOW_MO=700 npx playwright test --headed → watch it run in human time.
|
|
launchOptions: { slowMo: Number(process.env.SLOW_MO ?? 0) },
|
|
},
|
|
projects: [{ name: "chromium", use: { ...devices["Desktop Chrome"] } }],
|
|
webServer: {
|
|
// Real langgraph dev: real agent graph + real webhook routes + the harness
|
|
// http app (fake GitHub/Slack + mock UIs). Only the LLM is faked.
|
|
command:
|
|
"uv run langgraph dev --config tests/e2e/langgraph.e2e.json " +
|
|
`--port ${PORT} --no-browser --allow-blocking --no-reload`,
|
|
cwd: repoRoot,
|
|
url: `${baseURL}/mock/github/data`,
|
|
reuseExistingServer: !process.env.CI,
|
|
timeout: 180_000,
|
|
},
|
|
});
|