Commit graph

119 commits

Author SHA1 Message Date
seahaven-openswe[bot]
c2bc7720cd
feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678)

Ports four upstream commits that add opt-in LLM call routing through the
LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs,
no-agent-attribution, bun toolchain).

- #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings
  toggle, admin UI section, wired into make_model for all graph entrypoints
- #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over
  platform LANGSMITH_API_KEY
- #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host,
  SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware
- #1678 (73b7d1c0): fix OpenAI Responses reasoning replay —
  SanitizeOpenAIResponsesMiddleware, store/include config for encrypted
  reasoning content, reasoning_effort coercion for Chat Completions fallback

Refs #134

* fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity

- Downgrade logger.warning to logger.debug in gateway_overrides for
  not-routed providers and missing API key (Bedrock is the default
  provider in this fork, so these are expected steady states)
- Add Bedrock to the LLMGatewaySection route-toggle description so
  admins know it is not routed through the gateway
- Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with
  server.py and reviewer.py
- Restore the Bedrock region comment in model.py that explains the
  AWS_REGION / AWS_DEFAULT_REGION precedence

Refs #138

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-09 14:44:15 -04:00
Adam Moussa
b3b0274403
feat: Re-land deferred upstream features on modular webhooks (#80) (#128)
* fix(webhooks): fall back to vision model for Slack/Linear image threads

Re-land upstream #1626 onto the modular webhook structure. When a
Slack mention or Linear issue carries images but the resolved model is
text-only, fall back to a vision-capable model instead of dropping the
images. Re-points default_vision_model_pair at the fork's image-capable
models (Opus 4.8 default, else any supports_images model) rather than
upstream's openai:/anthropic: provider filter.

Refs #80, upstream #1626

* fix(slack): persist trace_message_ts so web-handoff updates the trace reply

Re-land upstream #1630 onto the modular structure. The first-mention
store_slack_run_mapping call did not pass trace_message_ts, so it was
never persisted (nothing to preserve from on first mention) and
_notify_slack_web_handoff always skipped the trace-reply update on web
handoff. Pass it through and cover it with a test.

Refs #80, upstream #1630

* feat(slack): include channel context in Slack prompts

Re-land upstream #1633 onto the modular structure. Fetch cached Slack
channel metadata once per event (_get_slack_channel_context) and thread
it through the docs-plz gate, repo resolution, and process_slack_mention
so prompts carry the channel name and a clearly-marked untrusted
channel description. Avoids duplicate conversations.info calls.

Refs #80, upstream #1633

* feat(tools): add slack_start_new_thread breakout tool

Re-land upstream #1638 onto the modular structure. Adds the
slack_start_new_thread tool (posts a top-level Slack message and
dispatches a fresh agent run for a broken-out task via the durable
dispatch_agent_run contract), wires it into the agent tool list and
tools/__init__, adds prompt guidance, and excludes it from plan mode so
it can't bypass the approval flow. Tool imports only live modules.

Refs #80, upstream #1638

* feat(plan): notify Slack on plan approval

Re-land upstream #1632 onto the modular structure. When a plan is
approved via the dashboard approve endpoint, post a thread reply to the
originating Slack thread noting the comment count and approver, after
the follow-up run is dispatched. Slack post failures never break
approval. Adapted to the fork's approve_plan (no plan_markdown read).

Refs #80, upstream #1632

* feat(plan): publish plans from sandbox files

Re-land upstream #1635 onto the modular structure, completing the
partially-ported change so dev is internally consistent. save_plan now
takes a plan_file_path, reads the agent-authored Markdown file from
/workspace/plans/ (validating extension/location/UTF-8/size) and
publishes it, instead of taking a plan_markdown string. Removes
write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can
author the plan file, updates enter_plan_mode/reject_plan guidance and
the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev).

Refs #80, upstream #1635

* fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs

INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop
revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/
Linear image URL could 302-redirect the fetch to an internal host / cloud
metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot
is_url_safe check. Route image fetches through the same per-hop resolve+pin+
revalidate loop the http_request tool uses, lifted into url_safety as the shared
request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token
on redirect so it can't be replayed to a redirect target.

SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at
DEBUG; multimodal logged them at INFO on every fetch. Log host-only.

Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened
its reach by no longer dropping images for text-only models. Fixing on the base
branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression
tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
ea1845d41d
fix(security): fence and neutralize untrusted Slack channel description in agent prompt
Hardens SLACK-PI-001 (sh-security-review). Slack channel topic/purpose is
editable by ordinary channel members and flowed verbatim into the agent LLM
prompt behind only a prose 'untrusted' label — an indirect prompt-injection
vector for an agent with network egress and repo write. Now strip leading
markdown structural tokens per line (so it can't forge the prompt's real
request/section delimiters), cap length, and wrap it in a per-render
unguessable sentinel fence (so injected text can't spoof a closing marker to
escape the data block). Deliberately diverges from upstream #1633.
2026-07-03 16:08:35 -04:00
seahaven-openswe[bot]
2b01652754
refactor: adopt modular webhook architecture (#1621) + port fork customizations (#85)
* Adopt upstream modular webhook skeleton (#1621)

Apply the durable-interrupt-dispatch refactor: split the monolithic
webapp.py into a thin routing layer plus per-source handlers in
webhooks/{github,slack,linear}.py, and add completion.py, dispatch.py,
and reconcile.py. Reconcile fork divergence by keeping the Bedrock/
Fireworks cross-provider fallback, the no-agent-attribution prompt
policy, the dashboard-handoff re-export, and the Slack channel-info
cache. ci_autofix is restored on the new dispatch model in a later
commit.

Refs: #80

* Port fork webhook security delta onto modular handlers

Re-apply the fork's security customizations that #1621 did not carry:
Linear webhook replay protection (freshness window on the signed
webhookTimestamp), per-repo token-cache binding threaded through the
thread token resolvers, the INTERNAL_BOT_LOGINS self-check in the
review-finding-reply path, and a user-mapping cache refresh before
email resolution on the issue and PR-comment paths (multi-replica
staleness). Existing fork security tests pass unchanged.

Refs: #80

* Restore CI auto-fix on the modular dispatch model

Bring back ci_autofix.py and the ci_monitor graph that #1621 deleted,
re-wiring the fork's security-reviewed PR-babysitting onto the new
structure: the CI-event, autofix-toggle, and review-feedback handlers
move into webhooks/github.py and the github_webhook router re-gains the
check_run/check_suite/workflow_run/status routing plus the autofix
command and actionable-review branches. Auto-fix runs now dispatch
through dispatch_agent_run (durability + completion webhook) while
keeping the deliberate batch-while-busy skip-rule via
get_thread_active_status. Restore langgraph.json's ci_monitor entry and
the fork autofix tests (dispatch mock + import paths re-pointed).

Refs: #80

* Reformat and update docs for the modular webhook split

Point CLAUDE.md and deploy/MIGRATION.md at the new webhooks/ modules
and the dispatch/completion/reconcile contract, and mark the
user-mapping cache-refresh fix as applied on the GitHub handlers.

Refs: #80

* Restore reject backstop for autofix dispatch

A burst of near-simultaneous CI events for one head SHA can slip past
the busy-check before the dedupe SHA is recorded, so dispatch the
autofix path with multitask_strategy=reject (dev's prior platform
default) to drop duplicate concurrent creates instead of letting them
interrupt each other. Also make the completion failure-reply dedup
claim-then-post and drop the unreachable interrupted branch.

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-06-30 18:46:46 -04:00
Adam Moussa
1f060f2a1d
chore: sync upstream/main, defer #1621 modular webhooks (#81)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
* chore: bake sfw binary into sandbox image (#1611)

sfw only ships a launcher that fetches its real binary at first run and does a
daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in
the sandbox (restricted egress; the proxy injects the GitHub App installation
token, which lacks access to that repo), so `sfw yarn install` errors with
"could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at
build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline.

* feat: editable plan mode + fix review-plan banner overlap (#1610)

* feat: editable plan mode + fix review-plan banner overlap

Lets the thread owner edit the plan markdown by hand from the plan-review
page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id}
endpoint that re-publishes the plan and mirrors it into the sandbox
plan.md, so approve hands the edited plan to the agent as the source of
truth. Also fixes the collapsed git-panel's floating expand button
covering the "Review plan ->" banner by reserving space for it.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: abort plan approval when the published plan read fails

get_plan_content() swallowed store errors and returned None, so a
transient failure during approve would still mark the plan approved and
dispatch the generic fallback text — silently dropping an owner's edited
plan. Read the plan strictly (raise_on_error=True) so approval aborts
instead, matching the comment read.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: show message timestamps (#1609)

* feat: show message timestamps

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: suppress fallback message timestamps

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: stable message + tool-call hover timestamps

Stamp a stable client-side arrival time per message and tool call (keyed
by id, persisted to localStorage). Messages render the timestamp inline;
tool rows reveal a dim timestamp chip on hover. Real backend created_at
still takes precedence when present.

* fix: hide client-stamped message timestamps

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add PR trace resolution (#1612)

* feat: add PR trace resolution

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: inject reviewer trace context as JSON

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review on PR trace resolution

Use the documented LangSmith metadata filter syntax
(and(eq(metadata_key,...), eq(metadata_value,...))) instead of
has(metadata, '{...}'), which does not match runs — _list_thread_runs
was silently returning nothing. Bound full-text searches to a 90-day
window so they don't hit LangSmith's large-window rate limit.

Also folds in the best-effort branch->head-sha resolver (dropping the
weighted scoring/threshold + repo/file evidence + GitHub hydration),
sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint.

The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session
were removed; resolution now runs deterministically from the trusted run
config with no model-controlled pr_url or thread_id.

* fix: scope branch trace search to the repo

Branch names like fix-tests aren't unique across repos (or older PRs) in
a shared tracing project, so an unscoped branch hit could resolve to an
unrelated thread and write its runs into the reviewer sandbox. Require
the repo slug to co-occur with the branch in matched runs; the full head
SHA stays unscoped since it is globally unique. Addresses open-swe review
on PR #1612.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: include plan links in PR descriptions (#1613)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: gate workflow pushes with approval (#1614)

* feat: gate workflow pushes with approval

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve proxy refresh test compatibility

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: bind workflow approvals to pushed ref

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: recover thread work as patch (#1615)

* feat: recover thread work as patch

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: search sandbox cwd for recovery patches

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: omit plan link in PR description when no plan exists (#1618)

Plan links in PR descriptions were always built from the thread id, so
runs that never produced a plan linked to an empty plan-review page.
Now the plan content store is consulted first; the link is only added
when a plan with non-empty markdown actually exists. A transient store
failure degrades gracefully (no link) rather than blocking PR creation.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add filter & grouping menu to agents threads sidebar (#1617)

Add a Cursor-style control to the agents sidebar that groups (None/Date/
Status/Project), filters (ownership, status, source, pull request, model,
repo, include-resolved), and compacts the threads list. All client-side over
already-fetched sidebar threads; preferences persist in localStorage.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: update langsmith sdk to 0.9.3 (#1616)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: clickable shared PR header in git panel and reviews (#1620)

* feat: clickable shared PR header in git panel and reviews

Replace the standalone "View PR" button in the agent git panel with a
clickable PR title, matching the reviews view. Extract a shared PrHeader
component reused by both the git panel and the review main body.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: drop PrHeader wrapper, use shared component directly

The review-side PrHeader was just a thin adapter mapping detail -> the
shared component's props. Inline it at the call site and use the shared
PrHeader directly so there's a single component.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: durable interrupt dispatch + completion webhook (#1621)

* wip(rebuild): core reliability spine

- remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring)
- dispatch core: agent/dispatch.py with multitask_strategy=interrupt +
  durability=sync + completion webhook; reroute all webhook + plan triggers;
  drop the racy in-process lock + is_thread_active busy-check
- completion webhook: agent/completion.py + /webhooks/run-complete loopback
  route for failure/timeout replies (idempotent)

Co-authored-by: open-swe[bot]

* feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning

Parallel batch on top of the reliability spine:
- async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the
  http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden
  the IP check to 'not is_global' (+ IPv4-mapped unwrap)
- reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list
  -> cancel_many), wired into the scheduler graph via task='reconcile'
- shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare
  httpx.AsyncClient() across utils/dashboard/webapp/middleware
- run budget: MODEL_CALL_RECURSION_LIMIT 5000->250
- fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8)
- drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls)
- confirm tool-result eviction + summarization auto-wired via backend
- slim system prompt ~8% (full harness-profile rewrite deferred)

Co-authored-by: open-swe[bot]

* feat(rebuild): harness-profile prompt + split webhooks out of webapp

- prompt.py: own the system prompt via a registered harness profile
  (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that
  share it stay safe), registered across all 4 providers; per-thread values
  stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k
  tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped
  ALL-CAPS markers.
- webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into
  agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the
  routes + tests; moved handlers reach shared helpers via the webapp namespace
  to preserve the test suite's monkeypatch targets.

Full suite: 1168 passing, lint clean.

Co-authored-by: open-swe[bot]

* Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks

Reverts the 250 cap from the run-budget change — long-running tasks legitimately
need many model calls. The notify_step_limit_reached safety net still fires if a
run does hit the cap, so runs end with a signal either way.

Co-authored-by: open-swe[bot]

* fix: address PR review (auth, SSRF, interrupted status, redirect headers)

- completion.py: drop `interrupted` from failure statuses — with
  multitask_strategy=interrupt a follow-up ends the prior run as interrupted,
  which is healthy, not a failure to report. [open-swe]
- /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when
  RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest.
  [corridor-security]
- SSRF: extract the URL validator to agent/utils/url_safety.py and apply it
  before server-side image fetches in multimodal.fetch_image_block.
  [corridor-security]
- http_request: preserve caller headers/extensions across redirect hops instead
  of dropping them on the first hop. [open-swe]

Co-authored-by: open-swe[bot]

* chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo)

Co-authored-by: open-swe[bot]

* fix: fail closed on run-complete webhook auth when secret unset

Corridor follow-up: verify_run_complete_token returns False (not True) when
RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never
unauthenticated. Logs a startup warning when the secret is absent, and dispatch
skips registering the webhook when there's no secret (no rejected callbacks).

Co-authored-by: open-swe[bot]

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: restore forced tool call to prevent premature run stops (#1622)

Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task.

Shipping to test whether it fixes runs that stop halfway through.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619)

Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1)

---
updated-dependencies:
- dependency-name: langgraph-checkpoint
  dependency-version: 4.1.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* fix: post reviewer resolution notes verbatim (#1624)

* fix: post reviewer resolution notes verbatim

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: stabilize dashboard follow-up e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve dashboard attribution in e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: make e2e attribution marker durable

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: only echo found e2e attribution

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: check live dashboard attribution in e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625)

Installs hung when prefixed with sfw inside the sandbox (trace 019f0608
stalled on a pending `sfw npm install` execute, never returned). Strip the
Socket Firewall guidance from the agent and reviewer prompts so installs run
through the project's package manager directly. sfw stays in the Docker image;
nothing invokes it now.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: make plan view mobile friendly (#1636)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: fall back to vision model for image threads (#1626)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: surface Slack thread errors (#1627)

* fix: surface Slack thread errors

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't set failure_reply_posted on Slack preprocessing errors

The preprocessing error handler was setting failure_reply_posted=True,
the same idempotency flag handle_run_completion checks to suppress
duplicate run-failure replies. Since preprocessing failures happen
before any run exists but the flag persists on the thread, a subsequent
run failure on the same thread would be silently ignored.

The preprocessing handler already posts its own Slack reply, so the
run-completion idempotency flag should not be set here.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: avoid recapping Slack replies (#1629)

* chore: avoid recapping Slack replies

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: simplify Slack reply prompt wording

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: update Slack trace reply on web handoff (#1630)

* fix: update Slack trace reply on web handoff

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: trigger web handoff on dashboard starts

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: format web handoff as contextual fragment

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve trace_message_ts when overwriting Slack run mapping

When store_slack_run_mapping is called without trace_message_ts (e.g. on
follow-up Slack mentions), it was unconditionally overwriting the
thread-level mapping and clobbering the timestamp captured from the
initial trace reply. After that, _notify_slack_web_handoff could not find
the original message, so a subsequent move to Web silently skipped the
Slack trace update.

Now, when trace_message_ts is not passed, the existing thread mapping is
read first and its trace_message_ts is preserved.

* style: ruff format

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643)

* fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures

shiki lazy-imports a grammar per language and these libs only live inside
lazy route components, so Vite's startup scanner never sees them. They get
discovered on first thread navigation, triggering a dep re-optimize +
force-reload that aborts the in-flight route-chunk import, surfacing as
"Failed to fetch dynamically imported module: .../$threadId.tsx".

Pre-bundle them (and the github themes + common code-block languages) via
optimizeDeps.include so the optimize happens once at startup. Dev-only;
production bundles are unaffected.

* fix: pre-bundle canonical shiki docker/make langs instead of aliases

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: show queued dashboard follow-ups (#1631)

* feat: show queued dashboard follow-ups

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: de-dupe queued follow-ups while streaming

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* feat: notify Slack on plan approval (#1632)

* feat: notify Slack on plan approval

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: post Slack approval notice after successful dispatch

Move the _maybe_post_plan_approved_to_slack call until after
_dispatch_followup succeeds so the Slack thread is not told
implementation is beginning before the LangGraph run is created.

Addresses PR review comment.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* feat: include Slack channel context in prompts (#1633)

Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>

* chore: keep plan guidance high-level (#1634)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: publish plans from sandbox files (#1635)

* feat: publish plans from sandbox files

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: avoid fixed plan filenames

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: virtualize local sandbox file paths

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve plan_file_path across set_plan_status

set_plan_status was rewriting the content record with only markdown
and status, dropping plan_file_path. After a reject, the owner's
dashboard edit would mirror to a different file than the agent's
original, and the next save_plan could republish the stale file.
Preserve plan_file_path when updating status.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: return to thread after plan approval (#1637)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add Slack breakout thread tool (#1638)

* feat: add Slack breakout thread tool

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: make fake LLM scripts declarative

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: exclude slack_start_new_thread from plan mode

The breakout tool can dispatch a fresh agent run that starts outside the
current plan-mode state, bypassing the approval flow. Add it to
PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating
tools while planning.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: require bun for ui agent work (#1639)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: request actions read for sandbox logs (#1642)

* fix: request actions read for sandbox logs

Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: restore actions:read scope after workflow push

After an approved workflow push, the guard was restoring the proxy with
BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read
scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which
includes actions: read) and fall back to BASE if the install hasn't
granted Actions read — mirroring the pattern in _create_sandbox_with_proxy.

Addresses review comment on PR #1642.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: widen split review diffs (#1647)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: install missing deps before verification (#1646)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: switch ui to pnpm (#1645)

* chore: require pnpm for ui agent work

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: switch ui to pnpm

Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* ci: use corepack for ui pnpm e2e build

Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add Sonnet 5 to model picker (#1651)

* chore: update Sonnet examples to Sonnet 5

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: add Sonnet 5 to model picker

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* Remove dead breakout-thread e2e scenario after dropping the tool

The merge resolution deferred upstream's Slack breakout-thread tool
(slack_start_new_thread, #1638) since it depends on the #1621 dispatch
module, but the e2e harness still scripted it. Removing the tool name
from fake_llm.py's _tool_step call left a malformed scenario, crashing
the langgraph-dev web server at import (TypeError: _tool_step() missing
'call_id') and failing Playwright E2E.

Drop the "breakout" script scenario, its _is_breakout_request helper +
ScriptRule, and the corresponding full_flow.spec.ts test.

* Revert upstream pnpm switch; keep bun for the UI build

The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/
global-setup.ts and ui/package.json, but our fork builds the UI with
bun (vercel.json + the E2E workflow's setup-bun). That left the
Playwright globalSetup running `corepack pnpm install --frozen-lockfile`
with no pnpm-lock.yaml, failing E2E at UI build time.

Revert global-setup.ts and ui/package.json to the dev (bun) baseline,
drop the merge-added ui/pnpm-lock.yaml, and remove the re-added
ui/AGENTS.md (our fork had deleted it).

* Align plan-review e2e + UI with the HEAD (pre-#1635) backend

The merge left a split plan vertical: the backend save_plan/plan_api are
HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/
#1637 per #80), but the plan UI and e2e harness were upstream's. The
fake_llm scenario called save_plan(plan_file_path=...) — upstream's
file-based #1635 contract — while HEAD save_plan takes plan_markdown,
so the plan never saved and PlanReview never rendered (E2E failure on
the plan-review locator).

Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts /
$threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the
whole plan flow (save -> render -> approve -> implement) is consistent
with the HEAD backend.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com>
Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
Adam Moussa
a30ce2ab40
feat: managed LangGraph Cloud + Vercel migration (PR2 — code fixes + docs) (#65)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint

Prepare the dashboard backend for the managed LangGraph Cloud + Vercel
runtime, where the API is HTTPS and cross-site from the UI.

- OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to
  https:// in _api_base_url() so GitHub stops rejecting login with
  "redirect_uri not associated with this application". _cookie_security()
  now treats a schemeless (managed) value as Secure; SameSite=None too,
  consistent with the coerced scheme.
- OAuth state cookie (#3): document that osw_oauth_state is host-only by
  design (a Domain cookie is unsafe across *.vercel.app, a public suffix),
  so login must always start on the stable alias to avoid "oauth state
  mismatch". Operational contract; no behavioral change.
- Admin user mappings (#4): add POST /admin/user-mappings so an admin can
  set the github_login -> work_email link from the dashboard instead of a
  raw Store write. New "admin" MappingSource provenance value.

* fix(webapp): refresh user-mapping cache on GitHub webhook paths

On managed LangGraph Cloud the backend runs multiple replicas, so the
per-process GitHub<->work-email mapping cache can be stale on the replica
handling a webhook (a mapping created on another replica is invisible
until refresh). process_github_pr_comment and process_github_issue now
refresh the cache from the durable Store before resolving the author's
email, matching the existing Slack mention path (process_slack_mention).

* perf(webapp): defer deepagents import to speed custom-app cold start

The custom FastAPI app (agent.webapp:app, the langgraph.json http.app)
pulled deepagents -> langchain_anthropic -> anthropic into its import
graph via dashboard.routes, only to build skill/chat seed files. Defer
those create_file_data imports into the functions that use them. Removes
deepagents/langchain_anthropic/anthropic from app import entirely and
roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger
cold-start saving since native anthropic init is skipped). Behavior
identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.)

* feat(ui): set work_email user mappings from the admin dashboard

Add an "Add / update" form to the admin User mappings section and the
adminUpsertUserMapping API client method, wiring the new
POST /admin/user-mappings endpoint. Admins can now create or update a
github_login -> work_email mapping directly instead of waiting for the
user to self-connect Slack.

* docs: document managed LangGraph Cloud + Vercel deployment

- INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL,
  DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty
  VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login
  and vercel.json stable-deployment-URL requirements, multi-replica cache
  note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting.
  Refresh the langgraph.json snippet to all six graphs.
- README: reframe deployment around the managed migration; link the plan.
- deploy/MIGRATION.md: import the self-hosted -> managed migration plan.
2026-06-29 19:58:38 -04:00
Adam Moussa
a4ed19ba61
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)

Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.

- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
  region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
  FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module

* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8

The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.

Map profile effort to additional_model_request_fields:
  {thinking: {type: adaptive, display: summarized},
   output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.

* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids

Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
  (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
  FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
  otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.

Surfaced by the cross-family review + verified against deploy/.

* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip

From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
  validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
  error code only, so the role ARN + account id in the raw botocore message never
  reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
  (Converse emits reasoning_content, not thinking) so the middleware is not a no-op
  on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)

* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids

Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
  us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
  each routed region (us-east-1/2, us-west-2). The model runs in the server process
  on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
  for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
  mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
  bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
  seed_store.sh's default via pick precedence, so the seed-script fix alone was
  insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
  ids to the Bedrock id (config.toml's model_id was an active, now-broken value).

AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.

* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)

Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
  (eval judge only — Bedrock builder/reviewer auth via the host IAM role).

REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00
Adam Moussa
8a9974c3c4
feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: default Slack/dashboard/schedule PRs + commits to the app identity (#57)

Slack/dashboard/schedule runs now author PRs and run git/gh operations as the
GitHub App seahaven-openswe[bot] by default (matching GitHub-issue runs), so the
self-review 422 is impossible by construction rather than guarded in the prompt.
A profile flag author_prs_as_user restores per-user attribution.

- open_pull_request._resolve_pr_author_token + auth.resolve_github_token: default
  to the installation token for these sources; per-user only when opted in.
- authorship: commit identity -> seahaven-openswe[bot] (numeric noreply;
  accepted Vercel-resolution risk, documented inline).
- self-trigger safety: INTERNAL_BOT_LOGINS + webapp/reviewer_reconcile/reply
  markers recognize seahaven-openswe[bot] (bot-authored events are now ours).

Supersedes the prompt-only guard in #58.

* fix: author commits as the app bot in the default path (SH-IDSPLIT-01)

Security review found the commit identity was NOT actually unified to the bot:
resolve_triggering_user_identity got a 403 from the installation token and fell
back to configurable['github_login'], so commits were still authored as the
triggering user (commit=user, push+PR=bot — a three-way split that missed the
stated goal). Now gate the triggering-user identity resolution on the same
default-bot decision as the token: slack/dashboard/schedule default to the app
bot identity unless author_prs_as_user is set.

* docs(security): record AUTHZ-SLACK-BOT-DEFAULT-001 as an accepted residual (#59)

Single-user deployment; bounded by App-on-pilot + ALLOWED_GITHUB_REPOS lock.
Revisit (add a per-user gate) before expanding users or the App installation.
2026-06-29 14:22:33 -04:00
Adam Moussa
a33aaec495
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54)
* fix: enforce a replay window on Linear webhooks (AUTHZ-001)

verify_linear_signature accepted any correctly-signed body with no freshness
check, so a captured request could be replayed indefinitely. Parse the
signed webhookTimestamp (Unix ms) and reject requests outside a 60s window,
failing closed when the field is missing or malformed — mirroring the Slack
verifier.

* fix: stop leaking upstream auth-error bodies into user comments

get_github_token_for_user folded the raw upstream response text into the
error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the
full body server-side only and return a generic "GitHub auth failed (status
<code>)". Also document the accepted shared-installation-token blast radius on
the bot-token-only path (AUTHZ-003).

* fix: bind sandbox and token caches to repo to prevent thread-id collision

A PR head-branch name is attacker-controllable and get_thread_id_from_branch
derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01).
The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on
thread_id alone, and a cached sandbox was reused after only an echo-ping, so a
different repo's webhook could bind to another thread's sandbox or token.

Without changing the persistent thread-id scheme:
- Persist the bound repo (owner/name) in thread metadata on sandbox creation and
  refuse to reuse a sandbox whose bound repo does not match the current event
  (SandboxRepoMismatchError); the in-memory proxy also carries the binding.
- Bind the GitHub-token cache entries to their repo and evict on a cross-repo
  read so a colliding thread_id cannot be served another repo's token.
- Thread repo through the reviewer and the webhook token resolvers.

* fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04)

The instance role and the GitHub deploy app role granted s3:ListBucket on the
whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts)
only ever lists under releases/, so add a StringLike s3:prefix=releases/*
condition. GetBucketLocation has no s3:prefix in its request context, so it
moves to its own unconditioned statement. Also document the accepted F-2
cross-env existence-oracle residual on BatchGetSecretValue.

* chore: suppress test-fixture credential false positive; document AUTHZ-002

Add a machine-level suppression for the fake Datadog key in the
test_team_credentials encryption-roundtrip fixture (CWE-798, not a real
credential). Clarify that the within-org thread-write path is intentional by
design (AUTHZ-002) — comment only, no behavior change.

* fix: casefold repo-binding keys to avoid spurious cross-repo mismatch

GitHub owner/name are case-insensitive. Casefold the owner/name key on both the
write (binding) and read (compare) sides — repo_cache_key and the metadata
bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate
same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2).

* fix: stop leaking upstream auth body in unexpected-result branch

The 2xx-but-missing-token/url branch echoed the parsed upstream response body
into the user-facing error. Return a generic message and log response_data
server-side only, mirroring the existing HTTPStatusError fix (Gap 4).

* fix: fail closed for unbound-legacy sandboxes and catch repo mismatch

Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no
recorded bound_repo (a pre-binding legacy thread, post-deploy) previously
reconnected-and-served the sandbox to the current repo, then rebound it. Now
fail closed: drop the stale id and recreate a fresh sandbox bound to this repo,
logging a reconnect-with-missing-binding event. A sandbox is never served to a
repo unless its binding is known and matches; new threads bind on first run
unchanged.

Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints,
log it for alarming, and surface a clean sanitized error instead of letting an
opaque deep-stack exception crash-loop the worker.

* chore: suppress test-fixture credential false positive in token-TTL tests

Add a machine-level suppression for the fake "ghp_secret" GitHub token used by
the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and
not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
Adam Moussa
f3db9f02e3
Adopt Sea Haven agent conventions, no attribution (#30)
Codify the box-only #4 customizations into Git so the AWS deployment
(which deploys from this repo) actually applies them — previously only
the retired sh-openswe box had them.

- prompt.py: branch names feature|bug|hotfix/<kebab> (optional <KEY->);
  imperative PR titles with no conventional-commit type: prefix; PR body
  Summary/Validation/Tests/Notes; handbook commit format. Rewrite the
  collaboration template from an attribution MANDATE to a PROHIBITION —
  no Co-authored-by bot trailer, no "Made by [Open SWE]" footer, no
  agent/AI notes on any artifact.
- github_comments.py: add @seahaven-openswe (the deployed App slug) to
  the mention triggers.
- authorship.py: remove the now-unused attribution helpers
  (build_pr_attribution_footer, add_bot_coauthor_trailer,
  add_pr_collaboration_note, PR_ATTRIBUTION_*). Keep OPEN_SWE_BOT_* —
  server.py still uses them for the sandbox git identity.
- Flip the attribution unit tests to assert the no-attribution behavior;
  drop tests for the removed helpers.

Commits stay authored as the triggering user for now — flipping
authorship to the bot account depends on the Vercel preview-deploy
constraint and is deferred to #11.

Refs: #4 #11

Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
2026-06-27 22:08:45 -04:00
Johannes du Plessis
4c7d45273f
perf: cut review-chat time-to-first-token (#1598)
* perf: cut review-chat time-to-first-token

The sandbox-less PR review chat paid several blocking network round-trips
before the first token on every message. Cache GitHub App installation
tokens in-process (per scope, until ~10m before expiry, above the proxy's
5m refresh window) so the chat graph factory and proxy stop re-minting one
each turn. Also drop the duplicate thread-metadata read in the commands
proxy and replace the heavy per-message get_review staleness check with a
single lightweight PR head-SHA lookup.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: keep review chat alive when reseed fails

Address review: a moved PR head now triggers _build_pr_context (and thus
get_review). For an existing chat, fall back to the last seeded context on
HTTPException instead of failing the command; fresh chats still surface the
error since they have no prior context.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:32:10 -07:00
Johannes du Plessis
3992d3ef5d
feat: repo-scoped dynamic sandbox snapshots (#1595)
* feat: repo-scoped dynamic sandbox snapshots

Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.

Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden repo snapshot builds

Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: document repo snapshot base image config

Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
Ramon Nogueira
3a0e2b4672
feat: plan mode with model-driven entry and collaborative review (#1580)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:06:58 -07:00
Brace Sproul
eac293d8a7
fix: serialize Slack run dispatch (#1591)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:04:08 -07:00
Johannes du Plessis
2df0eabd35
feat: reviewer enforces AGENTS.md/CLAUDE.md repo rules as mandatory pass (#1569)
* feat: reviewer enforces AGENTS.md/CLAUDE.md repo rules as mandatory pass

The reviewer already fetched AGENTS.md but treated violations as optional
candidate findings. Now the reviewer runs a dedicated compliance pass that
checks every changed hunk against each rule in AGENTS.md (or CLAUDE.md as
fallback), treating violations as mandatory findings rather than style nits.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: oversized AGENTS.md returns None instead of falling back to CLAUDE.md

Only a 404 (file absent) triggers fallback to CLAUDE.md. Oversize,
HTTP errors, and unexpected status codes now return None immediately
so the reviewer does not enforce stale rules from a secondary file.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 11:13:58 -07:00
Suraj Bayas
4030001ebf
feat: validate LLM API keys on startup (#1438)
* feat: validate LLM API keys on startup

* fix: correct relative import for options module

* refactor: move imports to top of file

* style: fix linting and formatting issues

* refactor: scope LLM validation to local dev and rename function

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
2026-06-18 09:24:46 -07:00
Johannes du Plessis
98b824bd54
feat: add size caps for PR diff, fetch_url, Slack threads, pagination, message queue [closes OPE-51] (#1567)
* feat: add size caps for PR diff, fetch_url, Slack threads, pagination, message queue

Per-source byte/token caps with explicit truncation markers to prevent
unbounded payloads from blowing up LLM context/memory.

- reviewer_diff.py: cap PR diff at 200K chars with head+tail truncation
- fetch_url.py: cap markdownify output at 100K chars
- slack.py: cap thread message fetch at 500 messages
- github_comments.py: cap _fetch_paginated at 50 pages
- thread_ops.py: cap queued messages at 100 (drop oldest)

Closes OPE-51

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: compute diff line set from full diff, keep most recent Slack messages

Address PR review comments:

1. Truncated diffs rejected valid findings: fetch_pr_diff now returns the
   full diff; truncate_diff is called separately in reviewer.py so the
   line set used for add_finding/publish_review validation is computed
   from the complete diff, not the truncated prompt text.

2. Slack cap dropped recent thread context: fetch_slack_thread_messages
   now keeps the most recent SLACK_THREAD_MAX_MESSAGES messages (was
   keeping the oldest). The tool surfaces a truncation marker in the
   formatted output so the LLM knows the thread was truncated.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 15:40:59 -07:00
Johannes du Plessis
7dd758f845
feat: shared GitHub HTTP helper with retries, rate-limit handling [closes OPE-45] (#1565)
* feat: shared GitHub HTTP helper with retries, rate-limit handling, and sane timeouts

Introduces agent/utils/github_http.py — a single place for GitHub API HTTP
calls with 30s/10s-connect timeouts (vs httpx's 5s default), exponential
backoff with jitter, Retry-After header support, and 429/secondary-rate-limit
detection. Migrates the reviewer publish path (reviewer_publish.py,
reviewer_diff.py, github_checks.py, github_ci.py) from one-shot
httpx.AsyncClient() calls to the shared helper.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't retry transport errors on non-idempotent GitHub writes

POST/DELETE/PATCH can create side effects server-side even when the client
gets a timeout or connection reset. Only retry transport errors for
idempotent methods (GET, HEAD, PUT, DELETE). 429/5xx status codes are still
retried for all methods since the server explicitly did not process the
request.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't retry 502/504 on non-idempotent GitHub writes

502 (bad gateway) and 504 (gateway timeout) are ambiguous — the upstream
may have processed the write before the gateway returned an error. Only
retry these for idempotent methods. 429 and 503 are still retried for all
methods since the server explicitly did not process the request.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 14:45:57 -07:00
Johannes du Plessis
b19804536c
feat: handle images sent to non-vision models in Slack, Linear, and web UI (#1560)
* feat: handle images sent to non-vision models in Slack, Linear, and web UI

Add vision capability checks across all image input paths. When a user
sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the
images are now skipped and a warning is injected into the prompt instead
of sending unsupported content to the model.

- Slack: resolve model at webhook time, skip image fetch + add warning
- Linear: same pattern as Slack
- Queued message middleware: read resolved model from thread metadata,
  strip images from queued payloads for text-only models
- Web UI: disable submit + show inline warning when images are attached
  to a non-vision model selection
- Shared: resolve_agent_model_id helper + vision_not_supported_warning

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: mock resolve_agent_model_id in Slack mention test

The test_process_slack_mention_queues_active_thread_message test was
missing a mock for the new resolve_agent_model_id call added to the
Slack webhook handler, causing a TypeError when image URLs triggered
the model resolution path.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: include vision warning in queued payload for text-only models

Update the prompt variable (not just content_blocks) before clearing
image_urls so the queued payload also carries the warning text when a
Slack/Linear follow-up arrives while the thread is busy.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:12:52 -07:00
Caroline di Vittorio
554d754585
feat: link PR attribution footer to the originating thread (#1539)
The "Made by Open SWE" PR footer linked to the generic dashboard
homepage. Point it at the dashboard thread that generated the PR
(/agents/<thread_id>), falling back to the homepage when no thread
id is available.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 09:50:18 -07:00
Johannes du Plessis
7397ff93ba
feat: CI auto-fix and PR babysitting for agent PRs (#1530)
* feat: CI auto-fix and PR babysitting for agent PRs

Watch CI failures and review feedback on PRs Open SWE opened, then dispatch
confidence-gated fix runs on the originating agent thread. Adds CI webhook
ingestion (check_run/check_suite/workflow_run/status), a per-PR @open-swe
autofix on|off toggle, auto-response to review comments, and a polling
ci_monitor graph that also flags merge conflicts. Gated by the existing
autofix_mode/trigger_mode settings, the enabled-repos opt-in, base-branch and
human-commit skip rules, dedupe, and a per-PR attempt cap.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review feedback on CI auto-fix

- Security: gate the no-mention review-feedback path on author trust —
  require a trusted author_association (OWNER/MEMBER/COLLABORATOR) plus a
  GitHub write/maintain/admin permission check before dispatching a
  write-capable agent run, preventing privilege escalation from
  read/triage/outside reviewers.
- Auth: reuse the originating PR thread's source + login/email when
  dispatching fix runs so the GitHub-token resolver authenticates them in
  non-bot-token deployments (bespoke github_ci source failed to resolve).
- Docs: document the Commit statuses: Read-only permission required for the
  Status webhook event.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 13:53:50 -07:00
Johannes du Plessis
889b2663d9
fix: point View trace link at the correct tracing project (#1521)
* fix: resolve trace URL project id by tracing project name

Graphs were split into separate LangSmith tracing projects
(open-swe-agent, open-swe-review) but the "View trace" link still used
a single fixed project-id env var pointing at the old combined project.
Resolve the project id from the tracing project name so agent and
reviewer links point at their respective projects, falling back to the
env var when resolution is unavailable.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* docs: document per-graph tracing projects for trace links

View trace links now resolve project IDs from the open-swe-agent /
open-swe-review project names. Document this so fresh deployments create
the right projects instead of relying on a single project ID.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-13 09:37:15 -07:00
Johannes du Plessis
bd5d5b24d4
fix: point reviewer "Open in Web" link to the review page (#1519)
The top-level review comment's "Open in Web" link pointed at the
agent thread (/agents/{thread_id}). Point it at the dashboard review
detail page (/agents/reviews/{owner}/{repo}/{number}) instead.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 14:15:58 -07:00
Johannes du Plessis
aa98d195a8
feat: route graphs to separate LangSmith tracing projects (#1508)
* feat: route graphs to separate LangSmith tracing projects

* docs: point graph entrypoint references at traced wrappers

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 17:57:16 -07:00
Johannes du Plessis
bfcb714c5c
feat: PR review page (#1495)
* feat: PR review page

* feat: manual review trigger card on admin page

* feat: finding focus mode with floating card and dimmed backdrop

* fix: dim only info panel, subtle block highlight for focused finding

* refactor: move PR reviews into agents page with file-tree sidebar

* fix: address review feedback

- paginate list_reviews until enough accessible records collected
- include rename/metadata-only files in diff API with empty hunks
- remount ReviewBody on head_sha change so viewed-files state resets

* fix: review tree background + finding card anchoring

* fix: remove no-op PR size chip and dead check links

* fix: parallel diff fetch, complete reReview type, shared github_headers

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 16:11:10 -07:00
Johannes du Plessis
194e9cd60e
feat: link Slack thread/Linear ticket in PRs and append ticket to title (#1504)
* feat: link Slack thread/Linear ticket in PRs and append ticket to title

When opening PRs, include any referenced Slack thread or Linear ticket
in the description and append the resolvable ticket number to the title.
Adds a Slack permalink to the webhook prompt so the agent has a
ready-to-link thread URL.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: use [closes <TICKET>] format in PR title suffix

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: deterministically append source refs to private-repo PRs only

Move the Slack/Linear source-reference linking out of the prompt and into
open_pull_request, gated to private repos so private Slack thread URLs and
Linear identifiers are never published to a public PR. Reverts the prompt
instruction and webhook permalink injection in favor of this server-side append.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 15:52:40 -07:00
Johannes du Plessis
ce3af2d9b2
fix: reviewer can silently review a stale checkout on reused sandboxes (#1503)
prepare_review_repo is best-effort, but the system prompt unconditionally
told the agent the repo is checked out at the PR head. On a reused sandbox,
a dirty worktree (or a transient fetch failure) makes 'git checkout <sha>'
fail; prep returns False, the old checkout stays in place, and the reviewer
confidently reads stale code — observed as 'I rechecked the current head'
replies quoting pre-push code.

- checkout with --force so leftover worktree state can't block it, verify
  HEAD matches the requested sha, tolerate 'git fetch --all' failures
- when prep fails, the prompt now warns the checkout may be stale and tells
  the agent to re-fetch/checkout (or fall back to API file contents) before
  trusting local files

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 13:42:10 -07:00
Johannes du Plessis
e0678e8c01
feat: track PR lifecycle state per thread for sidebar (#1492)
* feat: track PR lifecycle state per thread for sidebar

Persist a PR's draft/open/merged/closed state on the agent thread and keep
it in sync as PR webhooks fire, so the Agents UI can show per-thread PR
status the way Cursor does. open_pull_request now records the initial
draft/open state and pr_title; thread summaries expose diffStats; and the
PR webhook refreshes pr_state on close/reopen/draft/ready transitions.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: consolidate PR state mapping into shared derive_pr_state helper

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 12:21:59 -07:00
Johannes du Plessis
da20e57bcc
fix: refresh sandbox GitHub proxy token before mid-run expiry (#1496)
* fix: refresh sandbox GitHub proxy token before mid-run expiry

GitHub App installation tokens expire after exactly 1 hour. The LangSmith
sandbox proxy was configured once at run start with a snapshot of that
token, so runs longer than ~1h hit 401s on every gh/git call. Record the
proxy token's expiry per thread and add a before-model hook that
re-configures the proxy with a fresh token when it nears expiry.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve repo-scoped proxy token on mid-run refresh

Reviewer runs mint a repository-scoped installation token. Record the
repo scope per thread alongside the expiry so the before-model refresh
re-mints a token with the same scope instead of an installation-wide
token, avoiding privilege expansion on long reviewer runs.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: update passthrough stub for github_proxy_repositories param

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 10:59:21 -07:00
Christian Bromann
abf354bb05
feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475)
* feat(dashboard): stream agent chat via @langchain/react v2 protocol

Replace the bespoke SSE + React Query polling path with LangGraph’s
v2 event stream through credentialed dashboard proxies. Run starts go
through stream commands; mid-run follow-ups still queue via /messages.

* fix import path

* fix tests after rebase

* format

* PR feedback

* improved model fallback

* fix image handling

* embrace sdk

* cleanup

* cr

* more cleanup

* fix cors

* harden security

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 09:54:35 -07:00
Johannes du Plessis
259ec57186
fix: keep review check on follow-up commits, drop neutral conclusion (#1486)
* fix: keep review check visible on follow-up commits and stop using neutral

The "Open SWE Review" check completed as `neutral` whenever findings were
surfaced, which GitHub renders as a confusing "neutral check" group. Always
complete it as `success` (informational/non-blocking), matching Devin and
Corridor — the finding count stays in the title and findings post as comments.

Also create a fresh check run on the new head SHA in the push re-review path:
GitHub only shows check runs on a PR's current head, so the check vanished
after a follow-up push (and the stale id settled on an outdated commit).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: surface settled review check when push leaves diff unchanged

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 13:33:34 -07:00
Johannes du Plessis
e8770e1004
feat: report Open SWE Review as a PR check run (#1484)
* feat: report Open SWE Review as a PR check run

Auto-review dispatch now creates an in-progress 'Open SWE Review' check run
on the PR head SHA; publish_review completes it (neutral with findings,
success when clean). An after-agent hook fails the check if the run dies
before publishing. Requires the GitHub App's Checks: Read & write permission;
all calls are best-effort so a missing permission never breaks reviews.

* fix: address review feedback on check-run settling

Keep review_check_run_id when the completion PATCH fails so a later
publish or the after-agent hook can retry instead of hanging the check;
count out-of-diff findings toward the check conclusion.

* fix: retry failed check completion with the real publish conclusion

A transient PATCH failure after a successful publish previously left the
check id for the after-agent hook, which settled it as 'failure'. Persist
the intended result as review_check_pending_result and have the hook
prefer it over the generic failure fallback.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:53:48 -07:00
Johannes du Plessis
8f596a421d
feat: prep reviewer repo at init and load repo skills (#1480)
* feat: prep reviewer repo at init and load repo skills

Clone + checkout the PR head during reviewer agent init so SkillsMiddleware
can discover the repo's .agents/skills and .claude/skills from disk at its
one-shot scan, and so the LLM no longer narrates the clone mid-run.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: load reviewer skills from trusted base sha and fetch PR pull ref

Address review findings: skills are now extracted from the PR base sha via
git archive into a dir outside the checkout (prevents PR-authored SKILL.md
prompt injection), and repo prep fetches refs/pull/<n>/head with a strict
checkout so fork PRs fail loudly instead of silently reviewing the default
branch.

* fix: drop ref from skill-extraction log to satisfy CodeQL

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 09:54:32 -07:00
GowriH-1
85049290a3
feat: apply API standards skill in PR reviews for API changes (#1452)
Pull the api-standards skill from the LangSmith Context Hub at reviewer
run start and inject it into the system prompt, gated on the PR adding or
modifying an API surface. Best-effort: failures fall back to no supplement.

Co-authored-by: GowriH-1 <218394553+GowriH-1@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-08 14:37:04 -07:00
Mason Daugherty
ab2e0d1da3
feat: resolve Slack repo from channel topic/purpose (#1453)
Add conversations.info fetch so a repo:owner/name (or GitHub URL) token
in a Slack channel's topic/purpose pins the channel to a repo, slotting
in just below thread metadata in get_slack_repo_config.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-08 13:53:16 -07:00
Ramon Nogueira
d2268871eb
feat: surface dashboard UI link on PR reviews [INF-0000] (#1440)
* feat(reviewer): surface dashboard UI link on PR reviews [INF-0000]

Post a transient "review in progress" comment (with an "Open in Web"
dashboard link) when a reviewer run starts, then delete it once the
review lands. The published review body now carries the same
"Open in Web" link, so the link persists on the review itself.

The transient comment's id is tracked in reviewer thread metadata
(status_comment_id) so it can be deleted on completion.

* refactor(reviewer): inline dashboard URL helper, drop redundant future import [INF-0000]
2026-06-07 05:17:09 +00:00
Johannes du Plessis
3a91fbfabc
fix: revert co-author email to open-swe user noreply (#1424)
The bot's numeric noreply (215916821+open-swe[bot]@users.noreply.github.com)
introduced in ae946d1a doesn't resolve to a GitHub account Vercel accepts,
breaking preview deploys on PRs in langchainplus. Revert OPEN_SWE_BOT_EMAIL
to open-swe@users.noreply.github.com.

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 19:28:16 -07:00
Johannes du Plessis
6895ddcedc
feat: add scheduled web agents (#1422)
* feat: add scheduled web agents

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* fix: secure scheduled agent repositories

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* feat: rebuild scheduled agents as Automations tab

Scrap the inline ScheduledAgentsPanel and replace it with a dedicated
Automations tab: sidebar nav entry, list view with stat cards + empty
state, and a full editor (name, Active toggle, repo, scheduled trigger
picker, agent instructions + model).

* fix: clear collapsed-sidebar button on mobile in Automations

* fix: allow clearing automation repo on update

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-05 02:20:24 +00:00
Johannes du Plessis
f1baceb8ab
feat: add Gemini 3.5 Flash provider (#1420) 2026-06-04 17:38:54 -07:00
Johannes du Plessis
a26d5d32cb
feat: Add Fireworks model provider (#1415)
* feat: add Fireworks model provider with accurate reasoning levels

Adds Kimi K2.6, DeepSeek V4 Pro, Nemotron 3 Ultra, GLM 5.1 via fireworks: provider, mapping per-model reasoning_effort. Surfaces each model's real effort range (incl. openai none, deepseek xhigh/max).

* fix: drop serverless-unsupported Nemotron 3 Ultra, document FIREWORKS_API_KEY

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 21:59:47 +00:00
Johannes du Plessis
8b18e7e95d
fix: Reduce graph load and dashboard repo failures (#1412)
* fix: reduce graph load and dashboard repo failures

* fix: surface repo listing timeouts
2026-06-04 13:54:06 -07:00
Johannes du Plessis
60426e8b10
fix: reviewer resolves GitHub App token at run start, not via cross-process cache (#1409)
Reviewer/push runs execute in a worker process, but the GitHub token was
cached by the webhook handler in the API server process — a different
process — so the worker's in-process cache was always cold. resolve_github_token
then failed (User not authenticated / Unknown source: github_push) and only the
app-token fallback kept reviews working, noisily.

The reviewer always acts as the GitHub App (open-swe[bot]), so resolve the
installation token directly at run start, scoped to the repo. This also bypasses
org SAML enforcement that blocks user OAuth tokens. Drop the now-dead
cross-process cache writes in the webhook reviewer-dispatch handlers, and stop
leave_failure_comment raising on the github_push source.
2026-06-04 12:11:23 -07:00
Johannes du Plessis
c1e46b938e
feat: add Slack Block Kit reply options (#1407)
* feat: add Slack Block Kit reply options

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* docs: document Slack interactivity setup

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 10:26:09 -07:00
Brace Sproul
2070a770c2
fix: stop storing GitHub tokens in metadata (#1405)
* fix: stop storing GitHub tokens in metadata

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* fix: bound in-process GitHub token cache with 24h TTL + sweep

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
2026-06-04 09:33:51 -07:00
Johannes du Plessis
8913926dba
fix: remove manual review trigger surfaces (#1396)
Open SWE Review now runs from automated PR triggers, so drop the old Slack/GitHub review keyword entrypoints and keep PR comments on the regular agent path.
2026-06-03 12:33:22 -07:00
Mukil Loganathan
f4b27c0fee
feat: add Slack Open in Web link (#1392)
* feat: add Slack Open in Web link

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: skip web link for Slack reviewer runs

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: accept include_dashboard_link kwarg in Slack reviewer test double

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
2026-06-03 11:50:37 -07:00
open-swe[bot]
ae946d1aa5 chore: attribute commit co-author to open-swe[bot], not open-swe user
The Co-authored-by trailer and bot git identity used
open-swe@users.noreply.github.com, which resolves to the separate
open-swe *user* account rather than the open-swe[bot] GitHub App.
Switch OPEN_SWE_BOT_EMAIL to the bot's noreply address
(215916821+open-swe[bot]@users.noreply.github.com) so co-author credit
and the fallback author identity point at the bot.

Drive the prompt trailer and sandbox git config from the constant
instead of hardcoding the address.
2026-06-03 11:19:24 -07:00
Johannes du Plessis
0a2e682364
fix: scope public reviewer tokens (#1389)
* fix: scope public reviewer tokens

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: simplify reviewer token wiring; fix push re-scope + red test

- Remove the redundant _check_or_recreate_sandbox_for_proxy /
  _refresh_github_proxy_or_recreate_for_proxy wrappers and call the
  underlying functions directly (they already default the token to None).
- process_github_push_event: re-scope the GitHub App token when the push
  payload lacked repo privacy/id but PR metadata reveals a public repo, so
  reviewer.py never proxies a full-installation token for a public PR.
- Clarify the two-token sequence in trigger_pr_review_from_ref.
- Fix pre-existing failing test test_proxy_refresh_failure_recreates_sandbox
  and add coverage for _reviewer_token_for_repo + push-event scoping.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-03 09:25:17 -07:00
Johannes du Plessis
c8a997c125
fix: reliable, safe Slack account-connect prompt + first-login Slack dialog (#1383)
* fix: deliver Slack account-link prompt as a visible threaded reply

Blocked Slack users got no prompt at all. Prod logs show chat.postEphemeral
returns ok, but ephemeral messages are silently dropped in Slack's assistant
threads (where Open SWE runs), so the user sees nothing. Post the prompt as a
normal threaded reply instead — the same channel the agent uses to reply.

* fix: deliver Slack auth-failure prompt as a visible threaded reply

leave_failure_comment() tried an ephemeral message first and only fell back
to a thread reply on failure. Ephemeral messages succeed (ok) but are dropped
in Slack's assistant threads, so the fallback never fired and the user saw no
auth-failure prompt. Post the visible threaded reply directly, matching the
account-link prompt fix.

* fix: prompt blocked Slack users with a generic, token-free dashboard link

Addresses the review findings that posting the per-user account-link token /
auth URL in a visible thread lets any channel member bind their GitHub account
to the triggering user's Slack identity.

Drop the per-user signed link entirely. Both the account-link prompt
(_post_account_link_prompt) and the runtime auth-failure prompt
(leave_failure_comment) now post a plain dashboard settings link
(build_settings_url) as a visible threaded reply. The user signs in with GitHub
from their own session and connects Slack via verified OIDC on the settings
page — no secret in the thread, nothing to hijack, and no DM machinery.

* feat: nudge first-time users to connect Slack from the dashboard home

Show a Connect Slack banner on the agents landing page whenever Slack OAuth is
enabled and the user hasn't linked Slack yet. A first-time user (no Slack
mapping) sees it immediately after signing in; it disappears once connected.

* feat: prompt first-time users to connect Slack via a dialog

Replace the inline Connect Slack card on the agents home with a modal dialog
(Base UI). It opens automatically once the mapping query resolves to
"not connected" and closes itself once Slack is linked; "Maybe later" dismisses
it for the session. No new dependency — uses the design system's Base UI.

* copy: frame Slack connect as resolving the user's GitHub account

Drop 'act/reply on your behalf' wording across the connect-Slack dialog, the
Slack thread prompts (blocked + auth-failure), and the settings description.
Connecting Slack lets Open SWE resolve the user's GitHub account when they tag
it in Slack.
2026-06-02 20:55:07 -07:00
Johannes du Plessis
046388a4e9
feat: simplify PR attribution footer to "Made by Open SWE" (#1382)
Replace the double-attribution PR body footer (_Opened collaboratively by
{user} and open-swe._) with a single Cursor-style footer linking to the
project. Since PRs are now opened as the triggering user, the user no longer
needs to be named in the footer. The commit Co-authored-by trailer is kept.

Legacy footers are migrated on PR updates.
2026-06-02 19:52:27 -07:00