|
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: surface attributed PR creation failures Port upstream #1659: adds PullRequestCreationGuardMiddleware that blocks shell fallbacks (gh pr create, gh api /pulls, curl) when open_pull_request fails, keeping failures visible. Also adds preflight branch/repo visibility checks in open_pull_request with structured failure payloads, and updates the prompt to forbid PR creation fallbacks. Refs: #134 * fix: fall back to core GitHub App scope when optional grants missing (#1701) * fix: fall back to core GitHub App scope when optional grants missing Proxy-token minting requested workflows:write and actions:read in the permission set used for every sandbox. GitHub 422s a token request that asks for a permission the installation hasn't granted, so any install without workflows:write failed to mint a token and every run died in before-agent setup with "GitHub App installation token is unavailable". _resolve_proxy_token now walks a permission ladder (full -> +workflows -> core) and returns the first scope that mints, recording the granted scope so hourly proxy refreshes stay consistent. A missing optional grant now degrades to the install-time core scope instead of failing the run; workflow-file HITL pushes still require workflows:write and fail at push time when it is absent. * refactor: flatten proxy-token ladder loop with continue --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit f53caff1aa24a7b29d851b267aa3bdfe62c1e935) Sea Haven fork deviation: upstream #1701 folds workflows:write into the standing BASE/RUNTIME scope. This fork deliberately keeps workflows:write OUT of the standing permission ladder (RUNTIME = core + actions:read; LADDER = (RUNTIME, CORE)) so the sandbox proxy token cannot push .github/workflows/* during normal operation. workflows:write is minted only transiently by WorkflowPushGuardMiddleware for an approved HITL push and dropped on restore, preserving token scope as a backstop for the workflow- push approval control. Security-reviewed (agentic fan-out + GPT-4.1 cross review); the standing-scope-carries-workflows:write bypass was blocked. * fix(open-swe): harden proxy-token restore and mint error handling Two low-severity follow-ups from the security review of the #1701 port. Restore the recorded baseline scope after a workflow-push elevation instead of a hardcoded RUNTIME. An install granted workflows:write but not actions:read resolves its standing token to core; hardcoding RUNTIME on restore requested the ungranted actions:read, 422'd, and fired a false "SECURITY: failed to downscope" error on every approved workflow push before the core fallback recovered. The guard now captures the run's recorded scope before elevating (via the new get_recorded_proxy_permissions) and restores exactly that, falling back to the guaranteed core scope only when the baseline restore fails. Classify installation-token mint failures. get_github_app_installation_token_ with_expiry now treats HTTP 422 (a permission the installation hasn't granted) as the ladder's expected descend signal and keeps it at debug, while a non-422 failure (network/5xx/timeout) is surfaced at WARNING even when errors are otherwise suppressed — so a transient blip no longer silently downscopes a whole run under a debug-only trace. The reduced-scope warning no longer asserts a missing grant as the sole cause. * chore(triage): mark upstream #1701 landed on this branch Ported via PR #181 as Option A (workflows:write kept out of the standing proxy-token scope). Regenerated triage.md from triage.jsonl. * fix: restructure PR creation to POST-first with diagnose-on-failure Move preflight checks from an authoritative gate (before POST) to a diagnostic run after POST failure. This avoids false-positive failures when a just-pushed head branch is momentarily invisible to GitHub ref endpoints, and eliminates 2-3 extra serial API round-trips on the happy path. Also drop unused _PR_CREATED_FALSE indirection and add a docstring to pr_creation_guard acknowledging the fail-open detection design. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Adam Moussa <adam@seahavenind.com> |
||
|---|---|---|
| .. | ||
| screenshots | ||
| static | ||
| tests | ||
| .gitignore | ||
| agent_entrypoint.py | ||
| dev-mock.sh | ||
| e2e_env.py | ||
| fake_llm.py | ||
| fakes.py | ||
| global-setup.ts | ||
| harness.py | ||
| langgraph.e2e.json | ||
| package-lock.json | ||
| package.json | ||
| patches.py | ||
| playwright.config.ts | ||
| README.md | ||
Playwright E2E — the full Slack → implement → PR → reply flow
This drives the whole happy path through two mock UIs:
- A user asks Open SWE to implement something in a mock Slack thread.
- The real agent runs (via
langgraph dev): it implements the change in a local temp-dir sandbox, pushes a branch, and opens a PR on a fake GitHub. - It posts the PR link back to the same Slack thread — visible in the mock UI.
What is faked vs. real
Only the LLM and the external SaaS HTTP boundaries are faked. All agent code runs for real.
| Piece | Real or fake |
|---|---|
Slack webhook → process_slack_mention → run dispatch |
real (agent.webapp) |
get_agent, deepagents loop, tools, middleware, prompt |
real |
open_pull_request, slack_thread_reply tools |
real |
| Sandbox | real local provider, rooted in a throwaway temp dir |
| Git remote ("GitHub") | real git, a local bare repo the agent clones/pushes |
| The LLM | fake — a scripted model (fake_llm.py) emitting a fixed tool sequence |
api.github.com REST (PR create) + dashboard GitHub OAuth login |
fake (/fake-gh/...), state rendered at /mock/github |
slack.com/api (post message, etc.) |
fake (/fake-slack/...), thread rendered at /mock/slack |
GitHub App token mint, api.github.com/user identity |
stubbed (offline) |
The fake GitHub/Slack stores are the single source of truth the mock UIs render, so what Playwright asserts on is exactly what the real agent produced.
Files
e2e_env.py— env + constants set before anyagent.*import (sandbox=local, fake API URLs, isolatedGIT_CONFIG_GLOBAL, bot-token-only mode).fake_llm.py— the scriptedBaseChatModel(the only faked agent piece).patches.py— monkeypatches the boundaries (LLM, GitHub/Slack URLs, token mint).agent_entrypoint.py— langgraphagentgraph: applies patches, re-exports the realtraced_agent.harness.py— langgraphhttp.app: the realagent.webappplus the fake GitHub/Slack APIs, the mock UIs, and the control/compose endpoints.fakes.py— in-memory PR/Slack stores + git seeding of the bare remote.langgraph.e2e.json— dev-server config pointing at the two entrypoints above.static/{slack,github}.html— the mock Slack/GitHub UIs (external SaaS we can't run locally). The dashboard is not mocked — it's the realui/app.global-setup.ts— builds the realui/SPA (once) so the harness can serve it.
The dashboard — the real ui/ app
The dashboard is not mocked. The bot's "Open in Web" link
(DASHBOARD_BASE_URL/agents/{thread_id}) loads the actual built ui/ React
app — served same-origin from the harness so the session cookie and
/dashboard/api/* calls work without CORS. The signed session cookie is real
(minted via /control/login), so per-user authorization is genuine; the only
extra fake is the OAuth-token store (an external credential).
The UI is built by global-setup.ts with VITE_DASHBOARD_API_BASE_URL pointed at
the harness. It builds once; set E2E_FORCE_UI_BUILD=1 to rebuild (e.g. after a
UI change or port change). Requires Corepack with pnpm enabled.
Run
cd tests/e2e
npm install
npx playwright install chromium
npx playwright test # boots langgraph dev automatically, then runs
Watch it in human time:
SLOW_MO=700 npx playwright test --headed
Artifacts (replay a run)
Every test records a trace (DOM-snapshot timeline + network + console + source)
and a video; failures also get a screenshot. Locally they land in
test-results/<test>/ and are embedded in playwright-report/:
npx playwright show-report # browse runs; each has a Trace tab
npx playwright show-trace test-results/<test>/trace.zip # open one trace directly
In CI the Playwright E2E job uploads both playwright-report/ and
test-results/ as the playwright-report artifact on the run. Download it,
then npx playwright show-report <unzipped-dir> (or drag a trace.zip onto
https://trace.playwright.dev) to replay.
Poke at it by hand (from the repo root):
uv run langgraph dev --config tests/e2e/langgraph.e2e.json --port 2024 \
--no-browser --allow-blocking --no-reload
# open http://127.0.0.1:2024/mock/slack and /mock/github