* fix: fall back to core GitHub App scope when optional grants missing (#1701)
* fix: fall back to core GitHub App scope when optional grants missing
Proxy-token minting requested workflows:write and actions:read in the
permission set used for every sandbox. GitHub 422s a token request that
asks for a permission the installation hasn't granted, so any install
without workflows:write failed to mint a token and every run died in
before-agent setup with "GitHub App installation token is unavailable".
_resolve_proxy_token now walks a permission ladder (full -> +workflows ->
core) and returns the first scope that mints, recording the granted scope
so hourly proxy refreshes stay consistent. A missing optional grant now
degrades to the install-time core scope instead of failing the run;
workflow-file HITL pushes still require workflows:write and fail at push
time when it is absent.
* refactor: flatten proxy-token ladder loop with continue
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit f53caff1aa24a7b29d851b267aa3bdfe62c1e935)
Sea Haven fork deviation: upstream #1701 folds workflows:write into the
standing BASE/RUNTIME scope. This fork deliberately keeps workflows:write
OUT of the standing permission ladder (RUNTIME = core + actions:read;
LADDER = (RUNTIME, CORE)) so the sandbox proxy token cannot push
.github/workflows/* during normal operation. workflows:write is minted only
transiently by WorkflowPushGuardMiddleware for an approved HITL push and
dropped on restore, preserving token scope as a backstop for the workflow-
push approval control. Security-reviewed (agentic fan-out + GPT-4.1 cross
review); the standing-scope-carries-workflows:write bypass was blocked.
* fix(open-swe): harden proxy-token restore and mint error handling
Two low-severity follow-ups from the security review of the #1701 port.
Restore the recorded baseline scope after a workflow-push elevation instead
of a hardcoded RUNTIME. An install granted workflows:write but not actions:read
resolves its standing token to core; hardcoding RUNTIME on restore requested the
ungranted actions:read, 422'd, and fired a false "SECURITY: failed to downscope"
error on every approved workflow push before the core fallback recovered. The
guard now captures the run's recorded scope before elevating (via the new
get_recorded_proxy_permissions) and restores exactly that, falling back to the
guaranteed core scope only when the baseline restore fails.
Classify installation-token mint failures. get_github_app_installation_token_
with_expiry now treats HTTP 422 (a permission the installation hasn't granted)
as the ladder's expected descend signal and keeps it at debug, while a non-422
failure (network/5xx/timeout) is surfaced at WARNING even when errors are
otherwise suppressed — so a transient blip no longer silently downscopes a whole
run under a debug-only trace. The reduced-scope warning no longer asserts a
missing grant as the sole cause.
* chore(triage): mark upstream #1701 landed on this branch
Ported via PR #181 as Option A (workflows:write kept out of the standing
proxy-token scope). Regenerated triage.md from triage.jsonl.
---------
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Port upstream 304032fa: clarify that information-only answers should
check out relevant repos first for context, answer fully inline, and
only post a concise summary to Slack threads.
The fork already carried the shared-base Slack guidance; this adds the
missing TASK_EXECUTION_SECTION update and its test coverage.
Refs: #146
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
The WorkflowPushGuardMiddleware fails closed and blocks a git push whenever
it cannot parse the command into a plain, inspectable form (chained shell
operators, obfuscation, unrecognized refspecs). Most block reasons named only
the symptom ("chained shell commands"), so the agent kept retrying other
chained variants (cd &&, pushd &&) instead of dropping the chaining.
Append a single actionable remedy to every block reason, pointing at a plain
`git push origin <branch>` or `git -C <dir> push`, so a blocked run recovers
on the next attempt instead of looping.
* feat(models): re-add Fable 5 with admin disable toggle (port of upstream #1677)
* refactor(models): convert re-added Fable 5 to Bedrock model IDs
* fix(open-swe): correct Fable copy to describe provider data sharing, not ZDR
The ported admin toggle description and code comments described Fable 5 as
incompatible with Zero Data Retention. That is backwards: Fable 5 requires
the account to opt into Bedrock provider_data_share — prompts/completions are
retained and shared with Anthropic (up to 30 days, incl. human review). The
old UI copy would lead an admin to believe the opposite of what enabling the
toggle does. Reword the toggle description and the gate_fable_model /
team_settings comments accordingly. Still off by default. Refs #171.
* feat(dashboard): surface thread sandbox ID in sidebar
Port of upstream #1689 (feb7ac98): expose the thread's sandbox_id on the
dashboard thread summary and surface it in the sidebar via a "Copy sandbox
ID" action.
- thread_api: add sandboxId to the thread summary, hiding the
"__creating__" in-flight sentinel.
- queries/types: thread the sandboxId field through AgentThread and the
optimistic thread.
- AgentsSidebar: replace the hover-only resolve/delete buttons and the
right-click context menu with one touch-friendly kebab (⋮) menu that
works on both pointer and touch, and add the Copy sandbox ID item.
- vite: disable the PWA service worker in dev (it precaches assets and
defeats HMR).
- tests: unit test for the summary field + e2e spec in the real-backend
harness.
Closes#140. Part of #134.
* chore(triage): mark feb7ac98 (#1689) landed
Ported to dev via feat/thread-sandbox-id-sidebar (#140).
dev already pins fireworks-ai>=1.2.0a88 (uv.lock resolves a88) via #158
(efeb6b12), which jumped a85 -> a88 and exceeds the a86 that upstream
#1667 (fbc6de85) bumps to. Porting it would downgrade the constraint, so
mark the row landed/superseded rather than picking it.
Closes#141. Part of #134.
The built-in GITHUB_TOKEN cannot open the sync PR: the enterprise policy
blocks GitHub Actions from creating/approving pull requests, which
overrides the org and repo settings. That restriction applies only to
github-actions[bot], so mint a PROMOTE_APP installation token and pass
it to create-pull-request, mirroring promote-to-main.yml.
Add upstream-ledger-sync action to run daily to pull upstream commits, update triage.jsonl, and open a PR against dev so the list can be automatically updated to stay in line with upstream.
* feat: port plan-review & workflow-approval UX (#135)
Port six upstream commits onto dev:
- c03a6be7 (already ported): keep plan guidance high-level
- 546042a4: add workflow approval UI with diff preview, approval URLs,
web review links, and polling for approval status during active runs
- 216cf181: remove workflow token elevation; approved pushes pass
through directly without proxy token rewriting
- 3dbc0282: preserve plan redirects after login by accepting relative
same-origin redirect_to values and rejecting blocked paths
- bb104d93: submit plan comments with cmd+enter
- 90cb6caa: terse Slack replies, shared content via save_plan outside
plan mode (PLAN_STATUS_SHARED), reject shared-content mutations
Refs: #135
* feat: port durable dispatch hardening and startup latency improvements
Port five upstream PRs onto dev:
- #1621 / #1658: durable dispatch with loopback webhook defense,
create_durable_run helper, _config_with_prepare_run_id, degradation
to None for relative/loopback completion webhook URLs
- #1696: run-level completion webhook deduplication (replace
claim-then-post with post-then-flag per run_id), DeferredErrorModel
for graph-factory resilience, ToolRetryMiddleware for task subagents,
TimeoutWrapupMiddleware for all three graphs
- #1697: lazy-load __init__.py for agent.middleware, agent.tools,
agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search,
agent.webapp in request_pr_review, deepagents in sandbox.py); add
ttl_cache.py with stale-while-revalidate for tool loaders
Refs: #137
* fix: restore login page render and clear CI lint/format
The plan-review port removed the authRedirectUrl import from login.tsx
but left its call site, crashing the login page at runtime (blank page,
no 'Sign in to open-swe'). Pass the relative path straight to loginUrl,
matching the plan route and the backend relative-redirect handling.
Also drop an unused os import in the guard test and reformat
workflow_push_guard.py to satisfy ruff.
* feat: port model-fallback resilience from upstream (#1694, #1695)
- Add httpx.TransportError to the transient-exception set so
incomplete chunked reads on streamed responses trigger a
fallback instead of cancelling the run (#1694 / c9f6dd86).
- Rewrite fallback to alternate primary/fallback with exponential
backoff instead of a single failover, so the agent survives
multi-minute gateway outages spanning both providers (#1695 /
c9a9a7cd).
- Default backoff schedule (0, 5, 15, 30, 45) reaches past the
gateway's ~30s recovery window; jittered ±25%.
- On exhaustion, surface a terminal AIMessage explaining the
outage instead of crashing — progress is checkpointed so the
user can retrigger to continue.
- Preserve fork conventions: Bedrock ClientError retryability
check, sync wrap_model_call (using time.sleep instead of
asyncio.sleep), and existing access-error surfacing for
Anthropic, OpenAI, and Bedrock (botocore) provider errors.
- Update triage ledger (c9f6dd86, c9a9a7cd → landed) and
re-render triage.md.
Refs: #139
* fix: align workflow-push-guard tests with dev's transient-elevation impl
The dev merge auto-combined dev's elevation tests with the stale passthrough
tests inherited from the durable-dispatch branch; the passthrough tests
contradict dev's restored _run_with_workflow_token impl. Take dev's test file.
* fix: drop dead ttl_cache module; make fallback backoff jitter two-sided
ttl_cache.py was re-introduced via the dev merge but dev/#160 deliberately
removed it as dead code (no agent importer). Remove it to match dev.
Also make _jittered_delay symmetric (±25%) to match its docstring.
* chore(upstream-sync): triage 4 new upstream commits (#1708-#1713)
Synced ledger to upstream/main (71e3b818). New rows all deferred:
- #1708 add GPT-5.6 OpenAI models (FLAG-HUMAN: fork picker is Bedrock/Fireworks-only)
- #1709 stale admin model defaults after upgrades
- #1710 bump langchain-fireworks 1.4.4
- #1713 align reviewer eval with published findings
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
* feat: port plan-review & workflow-approval UX (#135)
Port six upstream commits onto dev:
- c03a6be7 (already ported): keep plan guidance high-level
- 546042a4: add workflow approval UI with diff preview, approval URLs,
web review links, and polling for approval status during active runs
- 216cf181: remove workflow token elevation; approved pushes pass
through directly without proxy token rewriting
- 3dbc0282: preserve plan redirects after login by accepting relative
same-origin redirect_to values and rejecting blocked paths
- bb104d93: submit plan comments with cmd+enter
- 90cb6caa: terse Slack replies, shared content via save_plan outside
plan mode (PLAN_STATUS_SHARED), reject shared-content mutations
Refs: #135
* feat: port durable dispatch hardening and startup latency improvements
Port five upstream PRs onto dev:
- #1621 / #1658: durable dispatch with loopback webhook defense,
create_durable_run helper, _config_with_prepare_run_id, degradation
to None for relative/loopback completion webhook URLs
- #1696: run-level completion webhook deduplication (replace
claim-then-post with post-then-flag per run_id), DeferredErrorModel
for graph-factory resilience, ToolRetryMiddleware for task subagents,
TimeoutWrapupMiddleware for all three graphs
- #1697: lazy-load __init__.py for agent.middleware, agent.tools,
agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search,
agent.webapp in request_pr_review, deepagents in sandbox.py); add
ttl_cache.py with stale-while-revalidate for tool loaders
Refs: #137
* fix: restore login page render and clear CI lint/format
The plan-review port removed the authRedirectUrl import from login.tsx
but left its call site, crashing the login page at runtime (blank page,
no 'Sign in to open-swe'). Pass the relative path straight to loginUrl,
matching the plan route and the backend relative-redirect handling.
Also drop an unused os import in the guard test and reformat
workflow_push_guard.py to satisfy ruff.
* fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary
* fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache
- Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__
(_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it.
- Reroute E2E model patching to deferred_model.make_model so make_model_or_defer
(used by all three graph factories) returns the scripted fake instead of
building a real model with fake credentials.
- Drop unused agent/utils/ttl_cache.py — no agent module imports it.
- Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001).
- Format tests/test_dispatch.py.
* fix: claim-then-post run-level failure dedup; stop permanent suppression
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
* feat: port plan-review & workflow-approval UX (#135)
Port six upstream commits onto dev:
- c03a6be7 (already ported): keep plan guidance high-level
- 546042a4: add workflow approval UI with diff preview, approval URLs,
web review links, and polling for approval status during active runs
- 216cf181: remove workflow token elevation; approved pushes pass
through directly without proxy token rewriting
- 3dbc0282: preserve plan redirects after login by accepting relative
same-origin redirect_to values and rejecting blocked paths
- bb104d93: submit plan comments with cmd+enter
- 90cb6caa: terse Slack replies, shared content via save_plan outside
plan mode (PLAN_STATUS_SHARED), reject shared-content mutations
Refs: #135
* fix: restore login page render and clear CI lint/format
The plan-review port removed the authRedirectUrl import from login.tsx
but left its call site, crashing the login page at runtime (blank page,
no 'Sign in to open-swe'). Pass the relative path straight to loginUrl,
matching the plan route and the backend relative-redirect handling.
Also drop an unused os import in the guard test and reformat
workflow_push_guard.py to satisfy ruff.
* fix: carry workflows:write on the standing proxy token
Complete the half-ported upstream 216cf181 cascade. The port dropped
_run_with_workflow_token from the guard but missed the paired github_app
change, so an approved .github/workflows push ran with the base token
(no workflows:write) and GitHub 403'd it.
Add workflows:write to BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS and delete the
now-orphaned WORKFLOW_RUNTIME_PROXY_TOKEN_PERMISSIONS constant; update the
github_app and proxy_auth tests to match. The HITL approval gate in
workflow_push_guard.py is unchanged — this only lets the standing token
push once a human approves.
* fix: restore transient workflow-token elevation (revert standing workflows:write)
The standing GitHub-App proxy token (BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS) is
ALWAYS-ON, so carrying workflows:write on it made the fork's HITL workflow-push
guard the sole control over unapproved workflow pushes. The guard's git-push
parser has gaps (obfuscated-expansion push, `gh api` REST contents PUT,
fully-qualified cross-branch refspecs); with a permanently workflows-scoped
token those gaps become live unapproved-workflow-push exploits (1 critical, 2
high — security review BLOCK on #159).
Restore dev's transient-elevation model:
- Drop workflows:write from BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS; re-add the
WORKFLOW_RUNTIME_PROXY_TOKEN_PERMISSIONS constant (base + workflows:write).
- Re-introduce _run_with_workflow_token in the guard: it mints the
workflows-scoped token via refresh_proxy_token around the approved,
guard-normalized fixed_command, then downscopes to RUNTIME then BASE in a
finally. Route the approval branch through it.
- Restore the dev token/elevation tests.
The standing token no longer carries workflows:write, so the three parser
bypasses hit GitHub 403 again; an approved push still succeeds because the
elevation grants workflows:write only around the normalized command. Keeps all
of #159's diff-preview / approval-URL / Slack-card guard additions.
* fix: reject protocol-relative path from sanitizeAuthRedirect (open redirect)
sanitizeAuthRedirect returned parsed.pathname+search+hash, which `new URL` can
resolve to a protocol-relative `//host` (e.g. input `/..//evil.com` normalizes
same-origin, passing the origin check, but yields a path starting with `//`).
ClientRedirect / login.tsx feed that path to window.location.replace, so it
navigates cross-origin — an open redirect. Reject any resolved path that is not
a single-leading-slash path (`^/[^/]`), falling back to the default. Adds
coverage for `/..//evil.com`, `/.//evil.com`, and `//evil.com`.
* fix: log SECURITY error when workflow-token downscope fails
The elevate->push->downscope finally block was silent on failure. If both
refresh_proxy_token calls fail, the sandbox retains workflows:write for the
rest of the run with no signal. Log a SECURITY error on the partial and full
downscope-failure paths so the retention is observable.
Addresses the GPT-4.1 cross-family review of the token-scope remediation.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
* feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678)
Ports four upstream commits that add opt-in LLM call routing through the
LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs,
no-agent-attribution, bun toolchain).
- #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings
toggle, admin UI section, wired into make_model for all graph entrypoints
- #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over
platform LANGSMITH_API_KEY
- #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host,
SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware
- #1678 (73b7d1c0): fix OpenAI Responses reasoning replay —
SanitizeOpenAIResponsesMiddleware, store/include config for encrypted
reasoning content, reasoning_effort coercion for Chat Completions fallback
Refs #134
* fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity
- Downgrade logger.warning to logger.debug in gateway_overrides for
not-routed providers and missing API key (Bedrock is the default
provider in this fork, so these are expected steady states)
- Add Bedrock to the LLMGatewaySection route-toggle description so
admins know it is not routed through the gateway
- Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with
server.py and reviewer.py
- Restore the Bedrock region comment in model.py that explains the
AWS_REGION / AWS_DEFAULT_REGION precedence
Refs #138
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
A sweep of the deferred backlog found two rows whose changes already shipped
in dev via #81 (the upstream-sync merge) — an empty cherry-pick confirms no net
change:
- #1618 (omit plan link in PR description when no plan exists) — was 'likely regression'
- #1646 (install missing deps before verification)
No code change; ledger accuracy only.
* feat: auto-load scoped AGENTS on reads (#1684)
Adapted from upstream langchain-ai/open-swe #1684 to the fork's
direct-import middleware registry and middleware stack ordering.
Adds SubdirAgentsReadMiddleware, which appends applicable ancestor
AGENTS.md instructions to read_file results once per run, so scoped
rules are visible before edits. Wired into get_agent immediately after
ToolErrorMiddleware, matching upstream's relative position.
Note: this changes file-read behavior for every main-agent run. The
reviewer graph uses its own leaner middleware stack and is unaffected.
(cherry picked from commit 7f7af71547be2199cea699284676f8ceefba7691)
* feat: add platform issue reporting tool (#1685)
Adapted from upstream langchain-ai/open-swe #1685 to the fork's
direct-import tool registry. Adds the report_platform_issue tool
(stdlib-only: returns a locally generated UUIDv7 report id, no external
network call) and wires it into get_agent's curated tool list.
Dropped upstream's test_task_retry_wraps_inside_tool_error_middleware
assertion, which references ToolRetryMiddleware that this fork does not
wire into the middleware stack.
(cherry picked from commit 88b62322b44103335773002d0745704cd96e9160)
---------
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
* feat: editable plan mode — owner hand-edits the plan before approval (#1610, #80)
Adds a PUT /dashboard/api/plan/{thread_id} endpoint and PlanReview UI edit
mode so the thread owner can refine the published plan markdown by hand. The
edited markdown is re-published as "ready" (preserving reviewer comments),
mirrored into the sandbox plan.md, and handed to the agent as the source of
truth on approve.
Approve now reads the published plan content strictly (raise_on_error=True)
so a transient store failure aborts instead of silently dropping an edited
plan, matching the comment-read contract.
The banner-overlap fix (collapsed git-panel clearing the "Review plan" link)
was already ported in #128; this picks up the remaining edit-mode pieces.
Refs #80
* fix(plan): make approve_plan idempotent, dispatch before persisting, fix comment count
SH-128-03: approve_plan set status APPROVED before dispatching the follow-up
run and had no already-approved guard, so a failed dispatch left the plan stuck
approved-but-undispatched and a double-submit double-dispatched + double-posted
the Slack notice. And the Slack notice counted len(comments) including empty
comments _format_comments filters out.
- Return 409 when the plan is already approved (idempotent double-click/retry).
- Dispatch the implementation run BEFORE persisting APPROVED so a dispatch
failure leaves the plan re-approvable. _dispatch_followup passes plan_mode
explicitly, so the run is unaffected by the reorder.
- Count only non-empty comments in the Slack approval notice.
Fixed here (not on #128) because #128's approve_plan is rewritten on this
branch; #129 inherits it. Adds tests for the 409, the filtered count, and the
dispatch-before-status ordering.
---------
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
* fix(webhooks): fall back to vision model for Slack/Linear image threads
Re-land upstream #1626 onto the modular webhook structure. When a
Slack mention or Linear issue carries images but the resolved model is
text-only, fall back to a vision-capable model instead of dropping the
images. Re-points default_vision_model_pair at the fork's image-capable
models (Opus 4.8 default, else any supports_images model) rather than
upstream's openai:/anthropic: provider filter.
Refs #80, upstream #1626
* fix(slack): persist trace_message_ts so web-handoff updates the trace reply
Re-land upstream #1630 onto the modular structure. The first-mention
store_slack_run_mapping call did not pass trace_message_ts, so it was
never persisted (nothing to preserve from on first mention) and
_notify_slack_web_handoff always skipped the trace-reply update on web
handoff. Pass it through and cover it with a test.
Refs #80, upstream #1630
* feat(slack): include channel context in Slack prompts
Re-land upstream #1633 onto the modular structure. Fetch cached Slack
channel metadata once per event (_get_slack_channel_context) and thread
it through the docs-plz gate, repo resolution, and process_slack_mention
so prompts carry the channel name and a clearly-marked untrusted
channel description. Avoids duplicate conversations.info calls.
Refs #80, upstream #1633
* feat(tools): add slack_start_new_thread breakout tool
Re-land upstream #1638 onto the modular structure. Adds the
slack_start_new_thread tool (posts a top-level Slack message and
dispatches a fresh agent run for a broken-out task via the durable
dispatch_agent_run contract), wires it into the agent tool list and
tools/__init__, adds prompt guidance, and excludes it from plan mode so
it can't bypass the approval flow. Tool imports only live modules.
Refs #80, upstream #1638
* feat(plan): notify Slack on plan approval
Re-land upstream #1632 onto the modular structure. When a plan is
approved via the dashboard approve endpoint, post a thread reply to the
originating Slack thread noting the comment count and approver, after
the follow-up run is dispatched. Slack post failures never break
approval. Adapted to the fork's approve_plan (no plan_markdown read).
Refs #80, upstream #1632
* feat(plan): publish plans from sandbox files
Re-land upstream #1635 onto the modular structure, completing the
partially-ported change so dev is internally consistent. save_plan now
takes a plan_file_path, reads the agent-authored Markdown file from
/workspace/plans/ (validating extension/location/UTF-8/size) and
publishes it, instead of taking a plan_markdown string. Removes
write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can
author the plan file, updates enter_plan_mode/reject_plan guidance and
the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev).
Refs #80, upstream #1635
* fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs
INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop
revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/
Linear image URL could 302-redirect the fetch to an internal host / cloud
metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot
is_url_safe check. Route image fetches through the same per-hop resolve+pin+
revalidate loop the http_request tool uses, lifted into url_safety as the shared
request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token
on redirect so it can't be replayed to a redirect target.
SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at
DEBUG; multimodal logged them at INFO on every fetch. Log host-only.
Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened
its reach by no longer dropping images for text-only models. Fixing on the base
branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression
tests (redirect-to-internal blocked; auth stripped on redirect).
* feat: add PR trace resolution (#1612)
* feat: add PR trace resolution
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: inject reviewer trace context as JSON
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on PR trace resolution
Use the documented LangSmith metadata filter syntax
(and(eq(metadata_key,...), eq(metadata_value,...))) instead of
has(metadata, '{...}'), which does not match runs — _list_thread_runs
was silently returning nothing. Bound full-text searches to a 90-day
window so they don't hit LangSmith's large-window rate limit.
Also folds in the best-effort branch->head-sha resolver (dropping the
weighted scoring/threshold + repo/file evidence + GitHub hydration),
sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint.
The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session
were removed; resolution now runs deterministically from the trusted run
config with no model-controlled pr_url or thread_id.
* fix: scope branch trace search to the repo
Branch names like fix-tests aren't unique across repos (or older PRs) in
a shared tracing project, so an unscoped branch hit could resolve to an
unrelated thread and write its runs into the reviewer sandbox. Require
the repo slug to co-occur with the branch in matched runs; the full head
SHA stays unscoped since it is globally unique. Addresses open-swe review
on PR #1612.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 69148f54f5)
* fix: post reviewer resolution notes verbatim (#1624)
* fix: post reviewer resolution notes verbatim
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: stabilize dashboard follow-up e2e
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve dashboard attribution in e2e
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: make e2e attribution marker durable
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: only echo found e2e attribution
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: check live dashboard attribution in e2e
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 5da3d0c657)
* chore: opt-in tracemalloc to attribute unclosed aiohttp sessions (#1657)
Prod logs show bursts of 'Unclosed client session' (aiohttp), leaking
fds + memory, but the warning omits the allocation site. When
DEBUG_TRACEMALLOC is set, start tracemalloc at webapp import so aiohttp
appends an 'Object allocated at' traceback naming the exact source.
Inert when the env var is unset.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 320bb39ab1f5cd7a2acdac52c7b375854334176c)
* feat: add PR review link route (#1698)
* feat: add PR review link route
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: avoid duplicate review shortcut runs
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 52fe29168814d7936ee6612b9c81f33339bb05b0)
* style: clean up leftover blank lines from cherry-pick conflict resolution
* fix(e2e): drop duplicate _ATTRIBUTION_RE from cherry-pick
The reviewer-misc pick re-added _ATTRIBUTION_RE next to _latest_attribution,
but the constant was already defined at module top (line 57, alongside
_PLAN_URL_RE) via the earlier #81 sync. Remove the redundant redefinition;
_latest_attribution resolves the surviving top-level constant.
---------
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
Co-authored-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
Hardens SLACK-PI-001 (sh-security-review). Slack channel topic/purpose is
editable by ordinary channel members and flowed verbatim into the agent LLM
prompt behind only a prose 'untrusted' label — an indirect prompt-injection
vector for an agent with network egress and repo write. Now strip leading
markdown structural tokens per line (so it can't forge the prompt's real
request/section delimiters), cap length, and wrap it in a per-render
unguessable sentinel fence (so injected text can't spoof a closing marker to
escape the data block). Deliberately diverges from upstream #1633.
* fix: surface Slack thread errors
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: don't set failure_reply_posted on Slack preprocessing errors
The preprocessing error handler was setting failure_reply_posted=True,
the same idempotency flag handle_run_completion checks to suppress
duplicate run-failure replies. Since preprocessing failures happen
before any run exists but the flag persists on the thread, a subsequent
run failure on the same thread would be silently ignored.
The preprocessing handler already posts its own Slack reply, so the
run-completion idempotency flag should not be set here.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit bb36448b0b)
* feat: add Slack breakout thread tool
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: make fake LLM scripts declarative
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: exclude slack_start_new_thread from plan mode
The breakout tool can dispatch a fresh agent run that starts outside the
current plan-mode state, bypassing the approval flow. Add it to
PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating
tools while planning.
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 747ce4bbe5)
Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
(cherry picked from commit 27d90ef196)
* chore(deps): bump vite-tsconfig-paths to ^6.1.1 and jsdom to ^29.1.1
Reconcile two Dependabot PRs (#107, #108) into one branch with a
single bun install so package.json and bun.lock stay consistent.
vite-tsconfig-paths: ^5.1.4 -> ^6.1.1 (dependencies)
jsdom: ^27.4.0 -> ^29.1.1 (devDependencies)
* Add --frozen-lockfile to CI Checks
Normal bun install treats bun.lock as updatable. If package.json requests a requirement bun.lock doesn't satisfy, bun quietly rewrites the lock and moves on. The committed lock is never updated.
* chore: Add trailing newline on new last run
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
Adds a triage-ledger job running `make triage-check` so a ledger edit that
forgets to regenerate the markdown (make triage-render) fails CI instead of
drifting silently. Stdlib-only, no deps.
#1651 added the family-aware `provider_fallback_pair` via `_claude_family_of`,
but the helper only matched `anthropic:claude-*` ids. This fork serves Claude
through Bedrock (`bedrock_converse:us.anthropic.claude-*`), so the family logic
was dead code: a dropped Bedrock Sonnet fell back to the Bedrock Opus that sits
first in the list instead of staying in the Sonnet family.
Teach `_claude_family_of` to parse `bedrock_converse` ids and add regression
coverage for the Sonnet-stays-on-Sonnet case.
* fix: make plan view mobile friendly (#1636)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 7ee3e05724)
* fix: return to thread after plan approval (#1637)
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit f32e492ab4)
* feat: reviews block agenda, sticky headers, accurate diff scroll (#1653)
Rework the AI-sorted blocks experience on the PR reviews page into a
Google-Docs-style outline: the left sidebar is now a clean number+title
agenda with scroll-spy highlighting of the active block; each block shows
its title + description (sticky) above its diff; and diff rows are pinned to
a uniform height so scroll-to lands precisely via the virtualizer's own
geometry instead of an estimate-driven correction loop.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 0b76afdc955e33805c7623d1502a75a9c7c9c1b7)
* fix: jump + ResizeObserver settle for review scroll-to (#1655)
Replace smooth-scroll plus frame-count correction loops on the PR
reviews page with an instant jump that re-asserts its target via a
ResizeObserver (the real "layout settled" signal). Block/file
navigation and finding/comment centering now land deterministically as
off-screen cards mount, files expand, and annotation cards measure,
instead of racing a smooth-scroll animation against height
reconciliation. Holds bail on user wheel/touch input and after a short
ceiling, and a new navigation cancels the previous hold.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
(cherry picked from commit 7530653bba7774d66a54b8bef0d2bbc25f519942)
* fix: purge expired thread_wakeup crons (#1656)
* fix: purge expired thread_wakeup crons
One-shot wakeup crons set an end_time that stops re-firing but the cron
row is never deleted, so dead rows accumulate (86 in prod). Add a purge
that deletes thread_wakeup crons past their end_time, called
opportunistically before scheduling a new wakeup, plus a one-time
backfill script. Conservative: matches only kind=thread_wakeup with a
past end_time.
* chore: retrigger Open SWE review
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 9e5a1924ef306269322c31342a1831e57831cfee)
* fix: add top padding to sticky review block header (#1660)
* fix: add top padding to sticky review block header
The sticky per-block header on the reviews page had padding below but
none above, so the block number badge sat glued against the top edge
when pinned. Add matching top padding for breathing room.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: use py-2 shorthand for review block header padding
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 23bd4a63fc5ba0fe853babf79ed33feb866cc8b2)
* fix: use global tokens for sidebar filter popover border (#1661)
The filter popover renders via base-ui Menu.Portal into document.body,
outside the .agents-ui container where the --ui-* CSS variables are
scoped. As a result border-[var(--ui-border)] resolved to an undefined
variable and border-color fell back to currentColor, producing a strong
near-black border (separators/hover/labels were similarly off).
Switch the portaled popup styling to the same global shadcn tokens the
theme/settings popover (SidebarUserMenu) already uses (border-border,
bg-border, bg-muted, text-muted-foreground). These are defined at :root
so they resolve inside portals too, and match the settings popover.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 63eb9a08209f683016abf01cdcc548bc5905f158)
* fix: preserve dashboard redirect after login (#1668)
* fix: preserve dashboard redirect after login
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: cover plan login redirect in e2e
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit bc7ce59169b5350da7286164afb83a7b037b528d)
* Disable React StrictMode (#1654)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 6575c327a3ac2b107a6e79a04fa61168d779dbf0)
* docs(upstream-sync): add cherry-pick runbook
Repo-specific runbook for bringing upstream (langchain-ai/open-swe) commits
into the fork: triage-sync discovery, the git cp workflow, the triage ledger,
themed-branch layout, and conflict/regression handling.
---------
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com>
* docs(upstream-sync): add triage ledger + cherry-pick hook plan
Seeds the upstream triage ledger (52 diverged commits from langchain-ai/open-swe
categorized: landed/won't-merge/deferred/untriaged) and the design plan for a
git-hook mechanism to keep it in sync during cherry-picks.
* feat(upstream-sync): jsonl-backed triage ledger + generator CLI
triage.jsonl is now the source of truth (52 rows migrated from triage.md);
triage.md is generated with a do-not-edit banner. scripts/triage.py provides
migrate/generate/reconcile/lookup/check-reject/set; make triage-render/check/reconcile
added (triage-check is CI-safe staleness gate). Stdlib-only so git hooks can call it.
* feat(upstream-sync): cherry-pick triage git hooks + git cp wrapper
post-commit journals each -x pick to an untracked .git-local journal; prepare-commit-msg
hard-blocks known-reject picks (commit-msg is a secondary backstop — clean picks skip it on
git 2.50.1), overridable via git cp --force / SH_CHERRYPICK_ALLOW_REJECT=1 /
sh.cherrypick.blockRejects=false. git-cp is the pre-apply guard + auto-reconcile. pre-push is
a SHIM that re-execs the global Sea Haven security pre-push so core.hooksPath=.githooks does
not shadow it; install-hooks.sh verifies that shim FIRST and refuses if it is missing.
* docs(upstream-sync): correct hook plan + git cp runbook
Record the verified git 2.50.1 finding that clean cherry-picks skip commit-msg, so the block
lives in prepare-commit-msg; note the locked HARD-BLOCK-by-default reject policy. CHERRYPICK.md
now leads with make install-hooks + git cp and explains the generated-md ledger.
* chore(upstream-sync): mark #1651 landed (Bedrock family fix on gateway-routing)
* docs(upstream-sync): rewrite CHERRYPICK.md as a repo-specific runbook
* feat(upstream-sync): add triage.py sync + make triage-sync
Discovers commits on dev..upstream/main not yet in the ledger and appends them
as untriaged (PR # and subject parsed from each commit), then bumps _meta
'last synced' to the upstream tip and regenerates triage.md. Closes the
discovery side of the workflow: triage-sync to pull in new work, git cp to land it.
* docs(upstream-sync): move cherry-pick runbook to PR1 branch as cherry-pick-runbook.md
Resolves Dependabot GHSA-9phm-9p8f-hw5m (open redirect via
protocol-relative URL in wildcard route rules) and GHSA-5w89-w975-hf9q
(proxy scope bypass via percent-encoded path traversal in routeRules).
nitro was pinned to "latest", which Dependabot can't resolve to a fixed
version, so the alerts stayed open even though the lockfile already
resolved 3.0.260603-beta (newer than the 3.0.260429-beta patch line).
Pin it to an exact version — nitro ships a date-stamped beta channel
where caret ranges behave unpredictably — so installs are reproducible
and both alerts close.
bun.lock also reconciles @pierre/trees beta.4 -> beta.5, which the
manifest already declared but the committed lockfile was stale on.
* Add Dependabot ignore for @types/node semver-major bumps
Prevent Dependabot from proposing wrong-direction @types/node major
bumps (e.g. 24 -> 26). /ui runs on Node 24 on Vercel; a too-new types
major still compiles but describes APIs absent at runtime.
Refs: #110
* feat(agent): seed all-repos custom instructions in default_prompt.md
Distill the universally-applicable Sea Haven authoring conventions into the
team-default Custom Instructions the main agent gets on every repo: secrets/
config placement, keep-docs-in-sync, verify-before-push, re-run-real-gates
after delegating, confirm-a-convention-before-adopting, and house writing
style. Toolchain references are generalized (not tied to a specific stack).
* feat(reviewer): seed Sea Haven review baseline as org-guidelines default
Bake DEFAULT_ORG_REVIEW_GUIDELINES (severity model, secrets, security surface,
tests, naming, deferred-work-needs-an-issue) and default org_guidelines to it
in _default_settings(). The reviewer now applies the Sea Haven baseline on
every repo until a workspace admin overrides it with a non-empty value via
the dashboard. Stack-agnostic and well under the 10k-char cap.
* refactor(prompt): consolidate duplicated COMMIT_PR_SECTION + add fork-sync runbook
COMMIT_PR_SECTION had two overlapping passes with a contradictory PR-title
rule (a fixed 'type: description' form vs the repo-aware detection). Collapse
into one numbered sequence (lint -> commit -> push/PR -> notify), keep the
authoritative repo-aware title rule, and drop the duplicate notify step. All
IMPORTANT directives (force-push ban, workflow-approval, autonomy, 403
handling) are preserved verbatim.
Add a fork-maintenance runbook to CLAUDE.md distilling the durable
upstream-sync methodology (conflict triage, deferred-refactor resolution rule,
the silent re-import/wiring hazards, test-impl-same-side, layered CI).
* refactor(prompt): adopt conventional-commit style
Flip the Sea Haven authoring convention baked into the agent prompt from
imperative/no-prefix to conventional-commit style:
- Commit subjects and the no-gate PR-title default now use
type(scope): description with the allowed type set (feat, fix, docs,
style, refactor, perf, test, build, ci, chore, revert, release).
- Branch prefixes expanded to feature/, fix/, hotfix/, chore/, docs/,
refactor/, release/ (kebab-case description).
- The repo-aware gate detection is preserved: a repo's own title gate
still wins and may narrow the allowed types/scopes.
Updated test_github_comment_prompts.py to assert the new convention.