open-swe/tests/e2e/playwright.config.ts

45 lines
1.7 KiB
TypeScript
Raw Permalink Normal View History

test(open-swe): add Playwright E2E for the Slack → PR → web handoff (#1583) * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
2026-06-22 15:54:46 -04:00
import { defineConfig, devices } from "@playwright/test";
import { resolve } from "node:path";
const repoRoot = resolve(__dirname, "..", "..");
const PORT = Number(process.env.E2E_PORT ?? 2024);
const baseURL = `http://127.0.0.1:${PORT}`;
export default defineConfig({
testDir: "./tests",
globalSetup: "./global-setup.ts",
fullyParallel: false,
workers: 1,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 1 : 0,
timeout: 90_000,
expect: { timeout: 60_000 },
reporter: [["list"], ["html", { open: "never" }]],
use: {
baseURL,
// Always capture the replayable artifacts: a trace (DOM snapshots, network,
// console, source — open with `npx playwright show-trace`) and a screen
// recording, plus a screenshot on failure. The CI job uploads them.
trace: "on",
video: "on",
screenshot: "only-on-failure",
// The built UI ships a PWA service worker; block it so tests never hit a
// stale cache and always see live API responses.
serviceWorkers: "block",
// SLOW_MO=700 npx playwright test --headed → watch it run in human time.
launchOptions: { slowMo: Number(process.env.SLOW_MO ?? 0) },
},
projects: [{ name: "chromium", use: { ...devices["Desktop Chrome"] } }],
webServer: {
// Real langgraph dev: real agent graph + real webhook routes + the harness
// http app (fake GitHub/Slack + mock UIs). Only the LLM is faked.
command:
"uv run langgraph dev --config tests/e2e/langgraph.e2e.json " +
`--port ${PORT} --no-browser --allow-blocking --no-reload`,
cwd: repoRoot,
url: `${baseURL}/mock/github/data`,
reuseExistingServer: !process.env.CI,
timeout: 180_000,
},
});