* fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect). |
||
|---|---|---|
| .. | ||
| screenshots | ||
| static | ||
| tests | ||
| .gitignore | ||
| agent_entrypoint.py | ||
| dev-mock.sh | ||
| e2e_env.py | ||
| fake_llm.py | ||
| fakes.py | ||
| global-setup.ts | ||
| harness.py | ||
| langgraph.e2e.json | ||
| package-lock.json | ||
| package.json | ||
| patches.py | ||
| playwright.config.ts | ||
| README.md | ||
Playwright E2E — the full Slack → implement → PR → reply flow
This drives the whole happy path through two mock UIs:
- A user asks Open SWE to implement something in a mock Slack thread.
- The real agent runs (via
langgraph dev): it implements the change in a local temp-dir sandbox, pushes a branch, and opens a PR on a fake GitHub. - It posts the PR link back to the same Slack thread — visible in the mock UI.
What is faked vs. real
Only the LLM and the external SaaS HTTP boundaries are faked. All agent code runs for real.
| Piece | Real or fake |
|---|---|
Slack webhook → process_slack_mention → run dispatch |
real (agent.webapp) |
get_agent, deepagents loop, tools, middleware, prompt |
real |
open_pull_request, slack_thread_reply tools |
real |
| Sandbox | real local provider, rooted in a throwaway temp dir |
| Git remote ("GitHub") | real git, a local bare repo the agent clones/pushes |
| The LLM | fake — a scripted model (fake_llm.py) emitting a fixed tool sequence |
api.github.com REST (PR create) + dashboard GitHub OAuth login |
fake (/fake-gh/...), state rendered at /mock/github |
slack.com/api (post message, etc.) |
fake (/fake-slack/...), thread rendered at /mock/slack |
GitHub App token mint, api.github.com/user identity |
stubbed (offline) |
The fake GitHub/Slack stores are the single source of truth the mock UIs render, so what Playwright asserts on is exactly what the real agent produced.
Files
e2e_env.py— env + constants set before anyagent.*import (sandbox=local, fake API URLs, isolatedGIT_CONFIG_GLOBAL, bot-token-only mode).fake_llm.py— the scriptedBaseChatModel(the only faked agent piece).patches.py— monkeypatches the boundaries (LLM, GitHub/Slack URLs, token mint).agent_entrypoint.py— langgraphagentgraph: applies patches, re-exports the realtraced_agent.harness.py— langgraphhttp.app: the realagent.webappplus the fake GitHub/Slack APIs, the mock UIs, and the control/compose endpoints.fakes.py— in-memory PR/Slack stores + git seeding of the bare remote.langgraph.e2e.json— dev-server config pointing at the two entrypoints above.static/{slack,github}.html— the mock Slack/GitHub UIs (external SaaS we can't run locally). The dashboard is not mocked — it's the realui/app.global-setup.ts— builds the realui/SPA (once) so the harness can serve it.
The dashboard — the real ui/ app
The dashboard is not mocked. The bot's "Open in Web" link
(DASHBOARD_BASE_URL/agents/{thread_id}) loads the actual built ui/ React
app — served same-origin from the harness so the session cookie and
/dashboard/api/* calls work without CORS. The signed session cookie is real
(minted via /control/login), so per-user authorization is genuine; the only
extra fake is the OAuth-token store (an external credential).
The UI is built by global-setup.ts with VITE_DASHBOARD_API_BASE_URL pointed at
the harness. It builds once; set E2E_FORCE_UI_BUILD=1 to rebuild (e.g. after a
UI change or port change). Requires Corepack with pnpm enabled.
Run
cd tests/e2e
npm install
npx playwright install chromium
npx playwright test # boots langgraph dev automatically, then runs
Watch it in human time:
SLOW_MO=700 npx playwright test --headed
Artifacts (replay a run)
Every test records a trace (DOM-snapshot timeline + network + console + source)
and a video; failures also get a screenshot. Locally they land in
test-results/<test>/ and are embedded in playwright-report/:
npx playwright show-report # browse runs; each has a Trace tab
npx playwright show-trace test-results/<test>/trace.zip # open one trace directly
In CI the Playwright E2E job uploads both playwright-report/ and
test-results/ as the playwright-report artifact on the run. Download it,
then npx playwright show-report <unzipped-dir> (or drag a trace.zip onto
https://trace.playwright.dev) to replay.
Poke at it by hand (from the repo root):
uv run langgraph dev --config tests/e2e/langgraph.e2e.json --port 2024 \
--no-browser --allow-blocking --no-reload
# open http://127.0.0.1:2024/mock/slack and /mock/github