Commit graph

40 commits

Author SHA1 Message Date
seahaven-openswe[bot]
68e02e58e0
Bump minor-level dependencies and fix FastAPI route test (#91)
Apply all minor-version bumps from the Dependabot group:
- fastapi 0.136.3 -> 0.138.2
- uvicorn 0.48.0 -> 0.49.0
- langchain-openai 1.2.2 -> 1.3.3
- exa-py 2.13.0 -> 2.15.0
- langchain-mcp-adapters 0.2.2 -> 0.3.0
- pytest 9.0.3 -> 9.1.1

FastAPI 0.137+ changed included routers to be stored as
_IncludedRouter wrappers instead of flattened routes. Update
test_plan_routes_registered to collect paths from the original
router so the plan routes remain discoverable.

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-06-30 19:46:27 -04:00
seahaven-openswe[bot]
901e66ba17
Bump patch-level dependencies from group update (#89)
Apply only the patch-version bumps from the minor-and-patch
group PR, leaving the minor bumps (fastapi, uvicorn,
langchain-openai, exa-py, langchain-mcp-adapters, pytest) on dev.

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-06-30 19:30:20 -04:00
Adam Moussa
1f060f2a1d
chore: sync upstream/main, defer #1621 modular webhooks (#81)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
* chore: bake sfw binary into sandbox image (#1611)

sfw only ships a launcher that fetches its real binary at first run and does a
daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in
the sandbox (restricted egress; the proxy injects the GitHub App installation
token, which lacks access to that repo), so `sfw yarn install` errors with
"could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at
build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline.

* feat: editable plan mode + fix review-plan banner overlap (#1610)

* feat: editable plan mode + fix review-plan banner overlap

Lets the thread owner edit the plan markdown by hand from the plan-review
page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id}
endpoint that re-publishes the plan and mirrors it into the sandbox
plan.md, so approve hands the edited plan to the agent as the source of
truth. Also fixes the collapsed git-panel's floating expand button
covering the "Review plan ->" banner by reserving space for it.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: abort plan approval when the published plan read fails

get_plan_content() swallowed store errors and returned None, so a
transient failure during approve would still mark the plan approved and
dispatch the generic fallback text — silently dropping an owner's edited
plan. Read the plan strictly (raise_on_error=True) so approval aborts
instead, matching the comment read.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: show message timestamps (#1609)

* feat: show message timestamps

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: suppress fallback message timestamps

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: stable message + tool-call hover timestamps

Stamp a stable client-side arrival time per message and tool call (keyed
by id, persisted to localStorage). Messages render the timestamp inline;
tool rows reveal a dim timestamp chip on hover. Real backend created_at
still takes precedence when present.

* fix: hide client-stamped message timestamps

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add PR trace resolution (#1612)

* feat: add PR trace resolution

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: inject reviewer trace context as JSON

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review on PR trace resolution

Use the documented LangSmith metadata filter syntax
(and(eq(metadata_key,...), eq(metadata_value,...))) instead of
has(metadata, '{...}'), which does not match runs — _list_thread_runs
was silently returning nothing. Bound full-text searches to a 90-day
window so they don't hit LangSmith's large-window rate limit.

Also folds in the best-effort branch->head-sha resolver (dropping the
weighted scoring/threshold + repo/file evidence + GitHub hydration),
sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint.

The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session
were removed; resolution now runs deterministically from the trusted run
config with no model-controlled pr_url or thread_id.

* fix: scope branch trace search to the repo

Branch names like fix-tests aren't unique across repos (or older PRs) in
a shared tracing project, so an unscoped branch hit could resolve to an
unrelated thread and write its runs into the reviewer sandbox. Require
the repo slug to co-occur with the branch in matched runs; the full head
SHA stays unscoped since it is globally unique. Addresses open-swe review
on PR #1612.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: include plan links in PR descriptions (#1613)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: gate workflow pushes with approval (#1614)

* feat: gate workflow pushes with approval

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve proxy refresh test compatibility

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: bind workflow approvals to pushed ref

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: recover thread work as patch (#1615)

* feat: recover thread work as patch

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: search sandbox cwd for recovery patches

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: omit plan link in PR description when no plan exists (#1618)

Plan links in PR descriptions were always built from the thread id, so
runs that never produced a plan linked to an empty plan-review page.
Now the plan content store is consulted first; the link is only added
when a plan with non-empty markdown actually exists. A transient store
failure degrades gracefully (no link) rather than blocking PR creation.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add filter & grouping menu to agents threads sidebar (#1617)

Add a Cursor-style control to the agents sidebar that groups (None/Date/
Status/Project), filters (ownership, status, source, pull request, model,
repo, include-resolved), and compacts the threads list. All client-side over
already-fetched sidebar threads; preferences persist in localStorage.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: update langsmith sdk to 0.9.3 (#1616)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: clickable shared PR header in git panel and reviews (#1620)

* feat: clickable shared PR header in git panel and reviews

Replace the standalone "View PR" button in the agent git panel with a
clickable PR title, matching the reviews view. Extract a shared PrHeader
component reused by both the git panel and the review main body.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: drop PrHeader wrapper, use shared component directly

The review-side PrHeader was just a thin adapter mapping detail -> the
shared component's props. Inline it at the call site and use the shared
PrHeader directly so there's a single component.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: durable interrupt dispatch + completion webhook (#1621)

* wip(rebuild): core reliability spine

- remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring)
- dispatch core: agent/dispatch.py with multitask_strategy=interrupt +
  durability=sync + completion webhook; reroute all webhook + plan triggers;
  drop the racy in-process lock + is_thread_active busy-check
- completion webhook: agent/completion.py + /webhooks/run-complete loopback
  route for failure/timeout replies (idempotent)

Co-authored-by: open-swe[bot]

* feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning

Parallel batch on top of the reliability spine:
- async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the
  http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden
  the IP check to 'not is_global' (+ IPv4-mapped unwrap)
- reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list
  -> cancel_many), wired into the scheduler graph via task='reconcile'
- shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare
  httpx.AsyncClient() across utils/dashboard/webapp/middleware
- run budget: MODEL_CALL_RECURSION_LIMIT 5000->250
- fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8)
- drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls)
- confirm tool-result eviction + summarization auto-wired via backend
- slim system prompt ~8% (full harness-profile rewrite deferred)

Co-authored-by: open-swe[bot]

* feat(rebuild): harness-profile prompt + split webhooks out of webapp

- prompt.py: own the system prompt via a registered harness profile
  (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that
  share it stay safe), registered across all 4 providers; per-thread values
  stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k
  tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped
  ALL-CAPS markers.
- webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into
  agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the
  routes + tests; moved handlers reach shared helpers via the webapp namespace
  to preserve the test suite's monkeypatch targets.

Full suite: 1168 passing, lint clean.

Co-authored-by: open-swe[bot]

* Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks

Reverts the 250 cap from the run-budget change — long-running tasks legitimately
need many model calls. The notify_step_limit_reached safety net still fires if a
run does hit the cap, so runs end with a signal either way.

Co-authored-by: open-swe[bot]

* fix: address PR review (auth, SSRF, interrupted status, redirect headers)

- completion.py: drop `interrupted` from failure statuses — with
  multitask_strategy=interrupt a follow-up ends the prior run as interrupted,
  which is healthy, not a failure to report. [open-swe]
- /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when
  RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest.
  [corridor-security]
- SSRF: extract the URL validator to agent/utils/url_safety.py and apply it
  before server-side image fetches in multimodal.fetch_image_block.
  [corridor-security]
- http_request: preserve caller headers/extensions across redirect hops instead
  of dropping them on the first hop. [open-swe]

Co-authored-by: open-swe[bot]

* chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo)

Co-authored-by: open-swe[bot]

* fix: fail closed on run-complete webhook auth when secret unset

Corridor follow-up: verify_run_complete_token returns False (not True) when
RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never
unauthenticated. Logs a startup warning when the secret is absent, and dispatch
skips registering the webhook when there's no secret (no rejected callbacks).

Co-authored-by: open-swe[bot]

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: restore forced tool call to prevent premature run stops (#1622)

Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task.

Shipping to test whether it fixes runs that stop halfway through.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619)

Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1)

---
updated-dependencies:
- dependency-name: langgraph-checkpoint
  dependency-version: 4.1.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* fix: post reviewer resolution notes verbatim (#1624)

* fix: post reviewer resolution notes verbatim

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: stabilize dashboard follow-up e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve dashboard attribution in e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: make e2e attribution marker durable

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: only echo found e2e attribution

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: check live dashboard attribution in e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625)

Installs hung when prefixed with sfw inside the sandbox (trace 019f0608
stalled on a pending `sfw npm install` execute, never returned). Strip the
Socket Firewall guidance from the agent and reviewer prompts so installs run
through the project's package manager directly. sfw stays in the Docker image;
nothing invokes it now.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: make plan view mobile friendly (#1636)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: fall back to vision model for image threads (#1626)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: surface Slack thread errors (#1627)

* fix: surface Slack thread errors

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't set failure_reply_posted on Slack preprocessing errors

The preprocessing error handler was setting failure_reply_posted=True,
the same idempotency flag handle_run_completion checks to suppress
duplicate run-failure replies. Since preprocessing failures happen
before any run exists but the flag persists on the thread, a subsequent
run failure on the same thread would be silently ignored.

The preprocessing handler already posts its own Slack reply, so the
run-completion idempotency flag should not be set here.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: avoid recapping Slack replies (#1629)

* chore: avoid recapping Slack replies

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: simplify Slack reply prompt wording

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: update Slack trace reply on web handoff (#1630)

* fix: update Slack trace reply on web handoff

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: trigger web handoff on dashboard starts

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: format web handoff as contextual fragment

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve trace_message_ts when overwriting Slack run mapping

When store_slack_run_mapping is called without trace_message_ts (e.g. on
follow-up Slack mentions), it was unconditionally overwriting the
thread-level mapping and clobbering the timestamp captured from the
initial trace reply. After that, _notify_slack_web_handoff could not find
the original message, so a subsequent move to Web silently skipped the
Slack trace update.

Now, when trace_message_ts is not passed, the existing thread mapping is
read first and its trace_message_ts is preserved.

* style: ruff format

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643)

* fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures

shiki lazy-imports a grammar per language and these libs only live inside
lazy route components, so Vite's startup scanner never sees them. They get
discovered on first thread navigation, triggering a dep re-optimize +
force-reload that aborts the in-flight route-chunk import, surfacing as
"Failed to fetch dynamically imported module: .../$threadId.tsx".

Pre-bundle them (and the github themes + common code-block languages) via
optimizeDeps.include so the optimize happens once at startup. Dev-only;
production bundles are unaffected.

* fix: pre-bundle canonical shiki docker/make langs instead of aliases

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: show queued dashboard follow-ups (#1631)

* feat: show queued dashboard follow-ups

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: de-dupe queued follow-ups while streaming

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* feat: notify Slack on plan approval (#1632)

* feat: notify Slack on plan approval

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: post Slack approval notice after successful dispatch

Move the _maybe_post_plan_approved_to_slack call until after
_dispatch_followup succeeds so the Slack thread is not told
implementation is beginning before the LangGraph run is created.

Addresses PR review comment.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* feat: include Slack channel context in prompts (#1633)

Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>

* chore: keep plan guidance high-level (#1634)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: publish plans from sandbox files (#1635)

* feat: publish plans from sandbox files

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: avoid fixed plan filenames

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: virtualize local sandbox file paths

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve plan_file_path across set_plan_status

set_plan_status was rewriting the content record with only markdown
and status, dropping plan_file_path. After a reject, the owner's
dashboard edit would mirror to a different file than the agent's
original, and the next save_plan could republish the stale file.
Preserve plan_file_path when updating status.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: return to thread after plan approval (#1637)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add Slack breakout thread tool (#1638)

* feat: add Slack breakout thread tool

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: make fake LLM scripts declarative

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: exclude slack_start_new_thread from plan mode

The breakout tool can dispatch a fresh agent run that starts outside the
current plan-mode state, bypassing the approval flow. Add it to
PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating
tools while planning.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: require bun for ui agent work (#1639)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: request actions read for sandbox logs (#1642)

* fix: request actions read for sandbox logs

Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: restore actions:read scope after workflow push

After an approved workflow push, the guard was restoring the proxy with
BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read
scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which
includes actions: read) and fall back to BASE if the install hasn't
granted Actions read — mirroring the pattern in _create_sandbox_with_proxy.

Addresses review comment on PR #1642.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: widen split review diffs (#1647)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: install missing deps before verification (#1646)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: switch ui to pnpm (#1645)

* chore: require pnpm for ui agent work

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: switch ui to pnpm

Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* ci: use corepack for ui pnpm e2e build

Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add Sonnet 5 to model picker (#1651)

* chore: update Sonnet examples to Sonnet 5

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: add Sonnet 5 to model picker

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* Remove dead breakout-thread e2e scenario after dropping the tool

The merge resolution deferred upstream's Slack breakout-thread tool
(slack_start_new_thread, #1638) since it depends on the #1621 dispatch
module, but the e2e harness still scripted it. Removing the tool name
from fake_llm.py's _tool_step call left a malformed scenario, crashing
the langgraph-dev web server at import (TypeError: _tool_step() missing
'call_id') and failing Playwright E2E.

Drop the "breakout" script scenario, its _is_breakout_request helper +
ScriptRule, and the corresponding full_flow.spec.ts test.

* Revert upstream pnpm switch; keep bun for the UI build

The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/
global-setup.ts and ui/package.json, but our fork builds the UI with
bun (vercel.json + the E2E workflow's setup-bun). That left the
Playwright globalSetup running `corepack pnpm install --frozen-lockfile`
with no pnpm-lock.yaml, failing E2E at UI build time.

Revert global-setup.ts and ui/package.json to the dev (bun) baseline,
drop the merge-added ui/pnpm-lock.yaml, and remove the re-added
ui/AGENTS.md (our fork had deleted it).

* Align plan-review e2e + UI with the HEAD (pre-#1635) backend

The merge left a split plan vertical: the backend save_plan/plan_api are
HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/
#1637 per #80), but the plan UI and e2e harness were upstream's. The
fake_llm scenario called save_plan(plan_file_path=...) — upstream's
file-based #1635 contract — while HEAD save_plan takes plan_markdown,
so the plan never saved and PlanReview never rendered (E2E failure on
the plan-review locator).

Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts /
$threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the
whole plan flow (save -> render -> approve -> implement) is consistent
with the HEAD backend.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com>
Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
dependabot[bot]
5caaf885a4
chore(deps): bump fireworks-ai from 1.2.0a75 to 1.2.0a85 (#37)
Bumps [fireworks-ai](https://fireworks.ai) from 1.2.0a75 to 1.2.0a85.

---
updated-dependencies:
- dependency-name: fireworks-ai
  dependency-version: 1.2.0a85
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-30 18:02:17 +00:00
Adam Moussa
a4ed19ba61
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)

Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.

- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
  region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
  FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module

* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8

The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.

Map profile effort to additional_model_request_fields:
  {thinking: {type: adaptive, display: summarized},
   output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.

* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids

Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
  (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
  FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
  otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.

Surfaced by the cross-family review + verified against deploy/.

* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip

From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
  validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
  error code only, so the role ARN + account id in the raw botocore message never
  reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
  (Converse emits reasoning_content, not thinking) so the middleware is not a no-op
  on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)

* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids

Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
  us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
  each routed region (us-east-1/2, us-west-2). The model runs in the server process
  on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
  for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
  mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
  bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
  seed_store.sh's default via pick precedence, so the seed-script fix alone was
  insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
  ids to the Bedrock id (config.toml's model_id was an active, now-broken value).

AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.

* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)

Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
  (eval judge only — Bedrock builder/reviewer auth via the host IAM role).

REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00
Adam Moussa
52cc696431
fix(deps): bump PyJWT to >=2.13.0 to close HS256 forgery (#47)
PyJWT <2.13.0 accepts a public-key JWK as an HMAC secret, letting an
attacker forge HS256 tokens when mixed key families are allowed
(GHSA high-sev Dependabot alert). 2.13.0 rejects the mismatch.

Direct dependency; uv.lock re-resolves 2.12.1 -> 2.13.0.
2026-06-28 20:41:27 -04:00
dependabot[bot]
2533402f6b
chore(deps): update langgraph-cli[inmem] requirement (#38)
Updates the requirements on [langgraph-cli[inmem]](https://github.com/langchain-ai/langgraph) to permit the latest version.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/cli==0.4.27...cli==0.4.30)

---
updated-dependencies:
- dependency-name: langgraph-cli[inmem]
  dependency-version: 0.4.30
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-28 02:40:34 +00:00
Ramon Nogueira
ca9280d25c
refactor(open-swe): plain HTTP comments instead of Yjs/BlockNote collab (#1601)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

* refactor(plan-mode): replace Yjs/BlockNote collab with plain HTTP comments

Drop the realtime collaborative editor (it can't work behind Vercel's
rewrite — WebSocket upgrades aren't proxied to the external LangGraph
backend) in favor of a simple whole-document comments API over plain HTTP.

Backend:
- Remove the Yjs WebSocket server (plan_collab.py), its lifespan, and the
  collab router; drop pycrdt / pycrdt-websocket deps.
- plan_store: replace the Yjs snapshot with comment CRUD (one store item per
  comment under ["plan","comments",thread_id]).
- plan_api: add GET/POST/DELETE comment endpoints; approve/reject now read
  comments server-side and format them for the follow-up run (no longer
  client-harvested). Comment delete is author-or-owner; approve stays owner-only.

Frontend:
- PlanReview renders the plan markdown read-only and shows a comments panel
  (list + add, polled every 4s for cross-user visibility).
- Drop @blocknote/*, y-websocket, yjs; lib/plan exposes get/add/deletePlanComment.

Tests: unit tests for the comments API + route registration; e2e drives the
HTTP comment UI (owner + collaborator, cross-user visibility, owner-only approve,
PR echoes the harvested feedback).

* fix(open-swe): clear stale plan comments on republish; fail loud on store errors

Address reviewer feedback:
- Clear comments when a revised plan is published (save_plan_content) so
  feedback on the prior revision doesn't resurface and get re-fed to the agent.
- list_plan_comments gains raise_on_error; approve/reject read comments before
  mutating state and propagate store failures (500) instead of silently
  dispatching the follow-up run with no feedback.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 22:12:42 +00:00
Ramon Nogueira
3a0e2b4672
feat: plan mode with model-driven entry and collaborative review (#1580)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:06:58 -07:00
dependabot[bot]
1451f3ff3e
chore(deps): bump langsmith from 0.8.11 to 0.8.18 (#1588)
Bumps [langsmith](https://github.com/langchain-ai/langsmith-sdk) from 0.8.11 to 0.8.18.
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/compare/v0.8.11...v0.8.18)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.8.18
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-23 11:16:49 -07:00
dependabot[bot]
71b9c9b631
chore(deps): bump langchain from 1.3.4 to 1.3.9 (#1551)
Bumps [langchain](https://github.com/langchain-ai/langchain) from 1.3.4 to 1.3.9.
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain==1.3.4...langchain==1.3.9)

---
updated-dependencies:
- dependency-name: langchain
  dependency-version: 1.3.9
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-16 19:39:02 -07:00
dependabot[bot]
82c805221e
chore(deps): bump langchain-anthropic from 1.4.4 to 1.4.6 (#1549)
Bumps [langchain-anthropic](https://github.com/langchain-ai/langchain) from 1.4.4 to 1.4.6.
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-anthropic==1.4.4...langchain-anthropic==1.4.6)

---
updated-dependencies:
- dependency-name: langchain-anthropic
  dependency-version: 1.4.6
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-16 15:41:09 -07:00
dependabot[bot]
85aaefc5bf
chore(deps): bump cryptography from 48.0.0 to 48.0.1 (#1550)
Bumps [cryptography](https://github.com/pyca/cryptography) from 48.0.0 to 48.0.1.
- [Changelog](https://github.com/pyca/cryptography/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pyca/cryptography/compare/48.0.0...48.0.1)

---
updated-dependencies:
- dependency-name: cryptography
  dependency-version: 48.0.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-16 15:40:44 -07:00
Johannes du Plessis
f539962c73
feat: server-side Datadog/LangSmith observability tools + team creds [closes OPE-54] (#1476)
* feat: server-side Datadog/LangSmith observability tools + team creds

Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review on observability tools

Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
  OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
  untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
  read-modify-write race dropping the other provider on concurrent saves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: async email resolution in observability authorization gate

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:07:42 -07:00
Johannes du Plessis
76f459ebe3
chore: bump langsmith to 0.8.11 and deepagents to 0.6.8 (#1464) 2026-06-09 10:42:19 -07:00
Johannes du Plessis
11042175c6
fix: roll back deepagents sandbox client (#1457)
Pin Deep Agents to the last pre-bump version so we can isolate lingering sandbox execute hangs after the LangSmith rollback.
2026-06-08 15:06:22 -07:00
Johannes du Plessis
ffb9366fc1
fix: roll back langsmith sandbox client (#1454)
Pin LangSmith to the last version verified in repo history before the websocket sandbox upgrade so production runs stop wedging while we investigate the SDK regression.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-08 13:42:07 -07:00
Johannes du Plessis
a26d5d32cb
feat: Add Fireworks model provider (#1415)
* feat: add Fireworks model provider with accurate reasoning levels

Adds Kimi K2.6, DeepSeek V4 Pro, Nemotron 3 Ultra, GLM 5.1 via fireworks: provider, mapping per-model reasoning_effort. Surfaces each model's real effort range (incl. openai none, deepseek xhigh/max).

* fix: drop serverless-unsupported Nemotron 3 Ultra, document FIREWORKS_API_KEY

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 21:59:47 +00:00
dependabot[bot]
a4404b711e
chore(deps): update langgraph-cli[inmem] requirement (#1387)
Updates the requirements on [langgraph-cli[inmem]](https://github.com/langchain-ai/langgraph) to permit the latest version.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/cli==0.4.24...cli==0.4.27)

---
updated-dependencies:
- dependency-name: langgraph-cli[inmem]
  dependency-version: 0.4.27
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-03 10:50:44 -07:00
dependabot[bot]
2cb617f5d9
chore(deps): bump cryptography in the major group across 1 directory (#1386)
Bumps the major group with 1 update in the / directory: [cryptography](https://github.com/pyca/cryptography).


Updates `cryptography` from 46.0.7 to 48.0.0
- [Changelog](https://github.com/pyca/cryptography/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pyca/cryptography/compare/46.0.7...48.0.0)

---
updated-dependencies:
- dependency-name: cryptography
  dependency-version: 48.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-03 10:50:33 -07:00
dependabot[bot]
b0d931406c
chore(deps): bump the minor-and-patch group with 13 updates (#1385)
Bumps the minor-and-patch group with 13 updates:

| Package | From | To |
| --- | --- | --- |
| [deepagents](https://github.com/langchain-ai/deepagents) | `0.6.6` | `0.6.7` |
| [fastapi](https://github.com/fastapi/fastapi) | `0.136.1` | `0.136.3` |
| [uvicorn](https://github.com/Kludex/uvicorn) | `0.46.0` | `0.48.0` |
| [langgraph-sdk](https://github.com/langchain-ai/langgraph) | `0.3.13` | `0.4.2` |
| [langchain](https://github.com/langchain-ai/langchain) | `1.3.2` | `1.3.4` |
| [langgraph](https://github.com/langchain-ai/langgraph) | `1.2.2` | `1.2.4` |
| [langchain-anthropic](https://github.com/langchain-ai/langchain) | `1.4.3` | `1.4.4` |
| [langsmith](https://github.com/langchain-ai/langsmith-sdk) | `0.8.3` | `0.8.8` |
| [langchain-openai](https://github.com/langchain-ai/langchain) | `1.2.1` | `1.2.2` |
| exa-py | `2.12.1` | `2.13.0` |
| [langchain-google-genai](https://github.com/langchain-ai/langchain-google) | `4.2.2` | `4.2.4` |
| [pytest-asyncio](https://github.com/pytest-dev/pytest-asyncio) | `1.3.0` | `1.4.0` |
| [ruff](https://github.com/astral-sh/ruff) | `0.15.12` | `0.15.15` |


Updates `deepagents` from 0.6.6 to 0.6.7
- [Release notes](https://github.com/langchain-ai/deepagents/releases)
- [Commits](https://github.com/langchain-ai/deepagents/compare/deepagents==0.6.6...deepagents==0.6.7)

Updates `fastapi` from 0.136.1 to 0.136.3
- [Release notes](https://github.com/fastapi/fastapi/releases)
- [Commits](https://github.com/fastapi/fastapi/compare/0.136.1...0.136.3)

Updates `uvicorn` from 0.46.0 to 0.48.0
- [Release notes](https://github.com/Kludex/uvicorn/releases)
- [Changelog](https://github.com/Kludex/uvicorn/blob/main/docs/release-notes.md)
- [Commits](https://github.com/Kludex/uvicorn/compare/0.46.0...0.48.0)

Updates `langgraph-sdk` from 0.3.13 to 0.4.2
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/0.3.13...0.4.2)

Updates `langchain` from 1.3.2 to 1.3.4
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain==1.3.2...langchain==1.3.4)

Updates `langgraph` from 1.2.2 to 1.2.4
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/1.2.2...1.2.4)

Updates `langchain-anthropic` from 1.4.3 to 1.4.4
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-anthropic==1.4.3...langchain-anthropic==1.4.4)

Updates `langsmith` from 0.8.3 to 0.8.8
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/compare/v0.8.3...v0.8.8)

Updates `langchain-openai` from 1.2.1 to 1.2.2
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-openai==1.2.1...langchain-openai==1.2.2)

Updates `exa-py` from 2.12.1 to 2.13.0

Updates `langchain-google-genai` from 4.2.2 to 4.2.4
- [Release notes](https://github.com/langchain-ai/langchain-google/releases)
- [Commits](https://github.com/langchain-ai/langchain-google/compare/libs/genai/v4.2.2...libs/genai/v4.2.4)

Updates `pytest-asyncio` from 1.3.0 to 1.4.0
- [Release notes](https://github.com/pytest-dev/pytest-asyncio/releases)
- [Commits](https://github.com/pytest-dev/pytest-asyncio/compare/v1.3.0...v1.4.0)

Updates `ruff` from 0.15.12 to 0.15.15
- [Release notes](https://github.com/astral-sh/ruff/releases)
- [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md)
- [Commits](https://github.com/astral-sh/ruff/compare/0.15.12...0.15.15)

---
updated-dependencies:
- dependency-name: deepagents
  dependency-version: 0.6.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: fastapi
  dependency-version: 0.136.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: uvicorn
  dependency-version: 0.48.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: minor-and-patch
- dependency-name: langgraph-sdk
  dependency-version: 0.4.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: minor-and-patch
- dependency-name: langchain
  dependency-version: 1.3.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: langgraph
  dependency-version: 1.2.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: langchain-anthropic
  dependency-version: 1.4.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: langsmith
  dependency-version: 0.8.8
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: langchain-openai
  dependency-version: 1.2.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: exa-py
  dependency-version: 2.13.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: minor-and-patch
- dependency-name: langchain-google-genai
  dependency-version: 4.2.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: pytest-asyncio
  dependency-version: 1.4.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: minor-and-patch
- dependency-name: ruff
  dependency-version: 0.15.15
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-03 10:23:09 -07:00
open-swe[bot]
8130a188ef
chore: bump deepagents to 0.6.6 (#1359)
* chore: bump deepagents to 0.6.6

Co-authored-by: Mason Daugherty <61371264+mdrxy@users.noreply.github.com>

* chore: remove obsolete deepagents reducer patch

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Mason Daugherty <61371264+mdrxy@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-29 11:08:55 -07:00
open-swe[bot]
aaeed1d95f
chore: bump deepagents to 0.6.4 (#1333)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Sydney Runkle <54324534+sydney-runkle@users.noreply.github.com>
2026-05-26 12:16:33 -07:00
open-swe[bot]
0ceca8d0db
chore: bump deepagents to 0.6.3 (#1322)
Pick up the patch release that restores message ID assignment for incoming messages without explicit IDs after the 0.6.2 regression.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Sydney Runkle <54324534+sydney-runkle@users.noreply.github.com>
2026-05-21 20:23:13 +00:00
Johannes du Plessis
82852f9eda
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt

Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.

Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
  defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
  or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
  claims without a concrete attacker/interleaving/scale, style preferences
  the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
  vendored / pure-rename hunks)

Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.

* trim prompt

* subagent prompting

* confidence ratings

* added medium

* enforce confidence threshold

* .

* reviewer: precision-tuned prompt + drop confidence gate

Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.

Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.

Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.

* benchmax

* adding google provider

* slight steering

* tuning

* more tuning

* fix

* cleanup

* reducing overfitting

* Add per-repo review style profiles and inject them into the reviewer.

Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix review style job errors leaking exception details to clients.

Return generic dashboard messages while logging full stack traces server-side.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 18:35:00 +00:00
open-swe[bot]
aca40ca397
chore: bump deepagents to 0.6.2 (#1316)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Sydney Runkle <54324534+sydney-runkle@users.noreply.github.com>
2026-05-20 15:33:31 +00:00
open-swe[bot]
c4d9ae4a67
feat: add idle TTL and delete-after-stop sandbox lifecycle controls (#1265)
* feat: add idle TTL and delete-after-stop sandbox lifecycle controls

* bump langsmith>=0.8.3, lower default idle TTL to 10 min

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-08 00:30:14 -04:00
open-swe[bot]
d2717ed33c
chore: bump deepagents to 0.5.7 (#1240)
Co-authored-by: Open SWE <open-swe@langchain.dev>
2026-05-05 12:02:08 -07:00
Johannes du Plessis
c6ea869df8
chore: bump all dependencies to latest stable versions (#1230)
Bump floors in pyproject.toml and refresh uv.lock. Notable changes:
- deepagents 0.5.3 -> 0.5.6
- langchain 1.2.15 -> 1.2.17, langchain-core -> 1.3.2
- langchain-anthropic 1.4.0 -> 1.4.2, anthropic -> 0.97.0
- langchain-openai 1.1.13 -> 1.2.1 (unpinned from ==)
- langgraph 1.1.6 -> 1.1.10, langgraph-api -> 0.8.5
- langsmith 0.7.32 -> 0.8.0
- fastapi 0.135.3 -> 0.136.1, starlette -> 1.0.0
- uvicorn 0.44.0 -> 0.46.0, pydantic -> 2.13.3

cryptography stays at <47 due to langgraph-api upper bound.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-01 13:43:34 -07:00
dependabot[bot]
213473dc08
chore(deps): bump the minor-and-patch group with 11 updates (#1198)
Bumps the minor-and-patch group with 11 updates:

| Package | From | To |
| --- | --- | --- |
| [deepagents](https://github.com/langchain-ai/deepagents) | `0.5.0a4` | `0.5.3` |
| [fastapi](https://github.com/fastapi/fastapi) | `0.128.3` | `0.135.3` |
| [uvicorn](https://github.com/Kludex/uvicorn) | `0.40.0` | `0.44.0` |
| [langgraph-sdk](https://github.com/langchain-ai/langgraph) | `0.3.4` | `0.3.13` |
| [langsmith](https://github.com/langchain-ai/langsmith-sdk) | `0.7.31` | `0.7.32` |
| [langchain-openai](https://github.com/langchain-ai/langchain) | `1.1.10` | `1.1.13` |
| [langchain-daytona](https://github.com/langchain-ai/deepagents) | `0.0.3` | `0.0.5` |
| [langchain-modal](https://github.com/langchain-ai/deepagents) | `0.0.2` | `0.0.3` |
| [langchain-runloop](https://github.com/langchain-ai/deepagents) | `0.0.3` | `0.0.4` |
| [exa-py](https://github.com/exa-labs/exa-py) | `2.10.1` | `2.12.0` |
| [ruff](https://github.com/astral-sh/ruff) | `0.15.0` | `0.15.10` |


Updates `deepagents` from 0.5.0a4 to 0.5.3
- [Release notes](https://github.com/langchain-ai/deepagents/releases)
- [Commits](https://github.com/langchain-ai/deepagents/compare/deepagents==0.5.0a4...deepagents==0.5.3)

Updates `fastapi` from 0.128.3 to 0.135.3
- [Release notes](https://github.com/fastapi/fastapi/releases)
- [Commits](https://github.com/fastapi/fastapi/compare/0.128.3...0.135.3)

Updates `uvicorn` from 0.40.0 to 0.44.0
- [Release notes](https://github.com/Kludex/uvicorn/releases)
- [Changelog](https://github.com/Kludex/uvicorn/blob/main/docs/release-notes.md)
- [Commits](https://github.com/Kludex/uvicorn/compare/0.40.0...0.44.0)

Updates `langgraph-sdk` from 0.3.4 to 0.3.13
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/0.3.4...0.3.13)

Updates `langsmith` from 0.7.31 to 0.7.32
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/compare/v0.7.31...v0.7.32)

Updates `langchain-openai` from 1.1.10 to 1.1.13
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-openai==1.1.10...langchain-openai==1.1.13)

Updates `langchain-daytona` from 0.0.3 to 0.0.5
- [Release notes](https://github.com/langchain-ai/deepagents/releases)
- [Commits](https://github.com/langchain-ai/deepagents/compare/langchain-daytona==0.0.3...langchain-daytona==0.0.5)

Updates `langchain-modal` from 0.0.2 to 0.0.3
- [Release notes](https://github.com/langchain-ai/deepagents/releases)
- [Commits](https://github.com/langchain-ai/deepagents/compare/langchain-modal==0.0.2...langchain-modal==0.0.3)

Updates `langchain-runloop` from 0.0.3 to 0.0.4
- [Release notes](https://github.com/langchain-ai/deepagents/releases)
- [Commits](https://github.com/langchain-ai/deepagents/compare/langchain-runloop==0.0.3...langchain-runloop==0.0.4)

Updates `exa-py` from 2.10.1 to 2.12.0
- [Commits](https://github.com/exa-labs/exa-py/commits)

Updates `ruff` from 0.15.0 to 0.15.10
- [Release notes](https://github.com/astral-sh/ruff/releases)
- [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md)
- [Commits](https://github.com/astral-sh/ruff/compare/0.15.0...0.15.10)

---
updated-dependencies:
- dependency-name: deepagents
  dependency-version: 0.5.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: fastapi
  dependency-version: 0.135.3
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: minor-and-patch
- dependency-name: uvicorn
  dependency-version: 0.44.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: minor-and-patch
- dependency-name: langgraph-sdk
  dependency-version: 0.3.13
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: langsmith
  dependency-version: 0.7.32
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: langchain-openai
  dependency-version: 1.1.13
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: langchain-daytona
  dependency-version: 0.0.5
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: langchain-modal
  dependency-version: 0.0.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: langchain-runloop
  dependency-version: 0.0.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
- dependency-name: exa-py
  dependency-version: 2.12.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: minor-and-patch
- dependency-name: ruff
  dependency-version: 0.15.10
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-16 07:09:02 +00:00
dependabot[bot]
aee6f20627
chore(deps): update langgraph-cli[inmem] requirement (#1199)
Updates the requirements on [langgraph-cli[inmem]](https://github.com/langchain-ai/langgraph) to permit the latest version.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/cli==0.4.12...cli==0.4.21)

---
updated-dependencies:
- dependency-name: langgraph-cli[inmem]
  dependency-version: 0.4.21
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-15 23:46:19 -07:00
dependabot[bot]
32eecbcee2
chore(deps): bump the uv group across 1 directory with 3 updates (#1191)
Bumps the uv group with 3 updates in the / directory: [langsmith](https://github.com/langchain-ai/langsmith-sdk), [pytest](https://github.com/pytest-dev/pytest) and [python-multipart](https://github.com/Kludex/python-multipart).


Updates `langsmith` from 0.7.25 to 0.7.31
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/compare/v0.7.25...v0.7.31)

Updates `pytest` from 9.0.2 to 9.0.3
- [Release notes](https://github.com/pytest-dev/pytest/releases)
- [Changelog](https://github.com/pytest-dev/pytest/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pytest-dev/pytest/compare/9.0.2...9.0.3)

Updates `python-multipart` from 0.0.22 to 0.0.26
- [Release notes](https://github.com/Kludex/python-multipart/releases)
- [Changelog](https://github.com/Kludex/python-multipart/blob/master/CHANGELOG.md)
- [Commits](https://github.com/Kludex/python-multipart/compare/0.0.22...0.0.26)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.7.31
  dependency-type: direct:production
  dependency-group: uv
- dependency-name: pytest
  dependency-version: 9.0.3
  dependency-type: direct:production
  dependency-group: uv
- dependency-name: python-multipart
  dependency-version: 0.0.26
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-15 23:31:23 -07:00
Aran Yogesh
4d4f5fbfc7
fix: proxy config restored the branch yogesh/GitHub auth proxy (#1173)
* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* removing logger.info

* formatting and linting

* fix: resolve lint errors in server.py (imports, unused vars, undefined names)

* feat: use opaque proxy headers for GitHub auth in sandbox

* linting formatting and test changes

* linting

* Delete .claude directory

* Delete tests/evals directory

* fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests

* fix: restore authorship, branch_name support, and installation token for PR creation

* linitng

* fix: move installation token fetch before commit, clean up dead proxy validation code

* fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config

* fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config

* linting

* linting

* fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox
2026-04-08 15:02:52 -07:00
Aran Yogesh
9aba4d0545
Revert "feat: authenticate git operations via sandbox proxy instead of creden…" (#1170)
This reverts commit 6305e13dc6.
2026-04-07 18:45:33 -07:00
Aran Yogesh
6305e13dc6
feat: authenticate git operations via sandbox proxy instead of credential files [closes: OPE-20] (#1070)
* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* removing logger.info

* formatting and linting

* fix: resolve lint errors in server.py (imports, unused vars, undefined names)

* feat: use opaque proxy headers for GitHub auth in sandbox

* linting formatting and test changes

* linting

* Delete .claude directory

* Delete tests/evals directory

* fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests

* fix: restore authorship, branch_name support, and installation token for PR creation

* linitng

* fix: move installation token fetch before commit, clean up dead proxy validation code

* fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config

* fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config

* linting

* linting

* fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox
2026-04-07 16:41:48 -07:00
open-swe[bot]
24a6343352
chore: upgrade deepagents to v0.5.0a4 (#1161)
Replace custom LangSmithBackend with built-in LangSmithSandbox from
deepagents. Update deprecated protocol method names in docs.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Sydney Runkle <54324534+sydney-runkle@users.noreply.github.com>
2026-04-06 10:24:50 -07:00
John Kennedy
d3506598fe
fix: upgrade deps to patch Pygments ReDoS and PyJWT crit header vulns (#1164)
- Bump deepagents >=0.4.3 → >=0.4.12
- Bump PyJWT >=2.8.0 → >=2.12.0 (fixes CVE-2026-32597, GHSA-752w-5fwx-jx9f)
- Pin Pygments >=2.20.0 in dev deps (fixes CVE-2026-4539, GHSA-5239-wwwm-4pmq)
- Regenerate uv.lock (also pulls langchain 1.2.15, langgraph 1.1.6)

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 20:42:51 -07:00
Aran Yogesh
7d1004ad66
fix: add Exa web search tool [closes: OPE-29] (#1133)
* fix: add Exa web search tool

* chore: update uv.lock for exa-py dependency

* linting

* chore: remove web_search from system prompt

* chore: drop search_type and category params from web_search
2026-03-25 16:24:31 -07:00
Brace Sproul
d7d9bc5179
fix: better custom backend support (#1071)
* fix: better custom backend support

* cr
2026-03-17 11:55:36 -07:00
Brace Sproul
bd52e5e09d
chore: Drop monorepo (#1029)
* chore: Drop monorepo

* cr
2026-03-06 16:10:34 -08:00
Renamed from apps/agent/pyproject.toml (Browse further)