Commit graph

282 commits

Author SHA1 Message Date
b2b5957233 fix(reviewer): harden verdict enforcement 2026-08-01 20:21:16 -04:00
eb91104db8 test(reviewer): cover team-enabled fork authorization 2026-08-01 20:21:16 -04:00
1b8f8d480d feat(reviewer): add dispatch verdict authorization 2026-08-01 20:21:16 -04:00
056688e822
fix(dashboard): authorize review verdicts 2026-07-31 12:34:56 -04:00
dependabot[bot]
aeec4f09dc
chore(deps-dev): bump @playwright/test
Bumps the minor-and-patch group with 1 update in the /tests/e2e directory: [@playwright/test](https://github.com/microsoft/playwright).


Updates `@playwright/test` from 1.61.1 to 1.62.1
- [Release notes](https://github.com/microsoft/playwright/releases)
- [Commits](https://github.com/microsoft/playwright/compare/v1.61.1...v1.62.1)

---
updated-dependencies:
- dependency-name: "@playwright/test"
  dependency-version: 1.62.1
  dependency-type: direct:development
  update-type: version-update:semver-minor
  dependency-group: minor-and-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-30 18:10:02 +00:00
Adam Moussa
d48cb12e08
feat(open-swe): port upstream clean batch (#1788, #1786, #1764, #1782, #1791, #1799) + guard hardening (#226)
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* Fix: Fix Insecure Direct Object Reference in slack_start_new_thread.py (#1788)

Co-authored-by: corridor-security[bot] <203152403+corridor-security[bot]@users.noreply.github.com>
(cherry picked from commit 32e81f2979a7baf11fe387df59f7d13a31889c74)

* Fix PR creation guard shell bypasses (#1786)

Co-authored-by: langsmith-fleet[bot] <langsmith-fleet[bot]@users.noreply.github.com>
(cherry picked from commit 75fb8b487852003916c4984504a13ee7226b2ceb)

* fix: add exc_info to swallowed exception in push re-review webhook (#1764)

(cherry picked from commit ab85b372b4f37b7feb849054553daed10852a42c)

* chore: clarify shared response image guidance (#1782)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 2e8ff4b72f1148bb36c0c1181063a3abd78b15d0)

* fix: match embedded review description background (#1791)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 4ea2441ada1229bc414b02950d821786be2f7301)

* fix: show current shared thread in sidebar (#1799)

* fix: show current shared thread in sidebar

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve resolved active sidebar threads

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
(cherry picked from commit a77c4e475643b4a55bb2f0c93c0aa2669014fbac)

* chore: switch deferred items to landed in upstream-sync triage documentation and jsonl entries (fork PR #226).

Signed-off-by: Adam Moussa <adam@seahavenind.com>

* harden PR guards + sidebar after security review

- Mirror upstream #1786's nested-shell / executable-normalization hardening
  into the fork-only pr_verdict_guard.py (verdict-gating is a real fork
  control), keeping it in parity with pr_creation_guard.py.
- Close the glued short-flag bypass (bash -c'...') in BOTH guards: a shell's
  -c argument can be concatenated into the same argv token, which the
  space-separated -c detection missed. Diverges pr_creation_guard.py from
  upstream #1786 by design; to be upstreamed.
- Gate the new #1799 sidebar active-thread refresh on ownership so a non-owner
  viewing a shared thread reads last-known state without persisting a metadata
  write (mirrors the is_owner gate on the single-thread read path).
- Fix an F821 in the #1799 cherry-pick (Mapping import / concrete dict type).

Guards remain intentionally fail-open per the honest-agent threat model;
docstrings narrowed to name the residual exotic-shell / stdin-fed vectors.

---------

Signed-off-by: Adam Moussa <adam@seahavenind.com>
Co-authored-by: corridor-security[bot] <203152403+corridor-security[bot]@users.noreply.github.com>
Co-authored-by: John Kennedy <65985482+jkennedyvz@users.noreply.github.com>
Co-authored-by: langsmith-fleet[bot] <langsmith-fleet[bot]@users.noreply.github.com>
Co-authored-by: Suraj Bayas <surajyou24@gmail.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-07-24 18:49:53 -04:00
Adam Moussa
c99bd78179
feat: clean-review auto-approve for unsolicited verdicts [BLOCKED — security] (#217)
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* feat(reviewer): clean-review auto-approve for unsolicited publish_review verdicts

An unsolicited publish_review(verdict="approve") — a run dispatched
without verdict_requested — is now honored when the review has zero open
findings, so clean auto-reviews land a real APPROVE. With open findings
it downgrades to a comment review (verdict_ignored_reason=
"approve_with_open_findings"). request_changes stays explicit-request-
only; the self-review, head-moved, and author-unknown downgrades and the
shell verdict guard are unchanged. The reviewer base prompt now instructs
the clean-approve call on auto-reviews.

* chore(security): record accepted-risk suppressions for clean-review auto-approve

Two confirmed-HIGH findings from /sh-security-review on the clean-review
auto-approve change are accepted and deferred (Adam, 2026-07-21), tracked
in #218. Machine-recorded per the mandatory-security-review policy; the
revisit trigger is promotion from dev to main/prod.
2026-07-21 20:18:25 -04:00
Adam Moussa
db05d5826f
feat(open-swe): port upstream clean batch (#1775, #1785, #1789) (#215)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* Fix: Fix Improper privilege management in server.py (#1789)

Co-authored-by: corridor-security[bot] <203152403+corridor-security[bot]@users.noreply.github.com>
(cherry picked from commit 3ea29d3f231dd66bd7627769b5659564be4525df)

* Fix: Fix Stored XSS in ReplyCard.tsx (#1785)

Co-authored-by: corridor-security[bot] <203152403+corridor-security[bot]@users.noreply.github.com>
(cherry picked from commit 31263f832a2ecedf669eee2e27b827a358a6a4d7)

* feat: add structured Linear issue filters (#1775)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit b5e529252c3012818289eca1b1cdeb8014721310)

* chore(triage): mark #1775 #1785 #1789 landed

Move the three clean cherry-picks in this batch from deferred to landed
in the upstream-sync ledger and re-render triage.md.

---------

Co-authored-by: corridor-security[bot] <203152403+corridor-security[bot]@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-07-20 19:54:16 +00:00
Adam Moussa
0f0f616cd4
feat(open-swe): explicit-request reviewer verdicts + shell verdict guard (#214)
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* feat(reviewer): explicit-request verdicts + shell verdict guard

Mention-triggered reviews that explicitly ask for a verdict now submit a
real APPROVE/REQUEST_CHANGES through publish_review; auto-reviews stay
advisory (COMMENT). Authorization is enforced in code: publish_review
honors a verdict only when the dispatching webhook set verdict_requested,
which only the explicit-mention path does.

- request_pr_review gains instructions (forwarded verbatim into an escaped
  requester_instructions data block) and request_verdict
- self-review guard downgrades verdicts on Open SWE-authored PRs; stale
  APPROVEs are best-effort dismissed when later findings land
- new PullRequestVerdictGuardMiddleware blocks gh pr review
  --approve/-a/--request-changes/-r, gh api, and curl verdict fallbacks on
  both the coding-agent and reviewer graphs
- shared escape helper moved to agent/utils/prompt_data.py

* fix(reviewer): harden verdict path against security-review findings

Adversarial security review (detector fan-out + proof-or-kill verifier)
of the verdict feature surfaced several verdict-integrity gaps; resolve
the confirmed ones:

- head-drift (high): a mid-run push moves the resolved head, so an APPROVE
  could anchor to an unreviewed commit. Downgrade any verdict to a comment
  when the resolved head differs from the reviewed head (verdict_ignored
  reason head_moved); the push's own re-review submits a fresh verdict.
- self-review fail-open: downgrade to comment when the PR author cannot be
  confirmed (author_unknown), and compare bot logins case-insensitively.
- verdict_submitted now reflects GitHub's returned review state, not just
  the event we asked for, so a coerced APPROVE isn't reported as submitted.
- an authorized verdict whose findings all anchor outside the diff now
  posts as a bodied review with zero inline comments instead of failing.
- add finding_reply to the shared data-block escape tag superset.
2026-07-20 15:28:00 -04:00
d980209157
test: cover bot-token fallback when an unbound user token is refused
Cross-family review (GPT-4.1) FIX item on the #1736 port: prove the
warn-and-drop path for unprincipaled user tokens leaves the cached bot
token reachable.
2026-07-17 17:08:09 -04:00
8d8d5bbbbf
fix: bind cached GitHub tokens to users (#1736)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 1ea0e600dcc234fa5a333c6f4b80b90e2e6679d3)

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-07-17 17:08:09 -04:00
a2794118ff
feat: optional separate LangSmith key/endpoint for sandboxes (#1760)
* feat: optional separate LangSmith key/endpoint for sandboxes

Adds optional SANDBOX_LANGSMITH_API_KEY / SANDBOX_LANGSMITH_ENDPOINT env
overrides so sandboxes can run against a different LangSmith workspace than
the one used for tracing and other API calls. Both fall back to the existing
LANGSMITH_API_KEY / LANGSMITH_ENDPOINT resolution, so default behavior is
unchanged.

Applied to sandbox create/connect/delete, the GitHub proxy config, and repo
snapshot builds.

* feat: name langsmith sandboxes openswe-<b32(thread id)>

New sandboxes get a deterministic, thread-traceable name derived from the
LangGraph thread id (UUID base32-encoded lowercase, no padding), e.g.
openswe-ci2fm6asgrlhqerukz4bencwpa. Falls back to an unset name when no thread
id is present. Reconnect/delete still key off the server-assigned sandbox id.

* fix: pass sandbox base URL (root + /v2/sandboxes) to langsmith SDK clients

The SDK's api_endpoint is the sandbox base, not the API root — its methods
append /boxes, /snapshots, etc. Passing the bare root sent calls to
<root>/boxes instead of <root>/v2/sandboxes/boxes. Add _get_sandbox_api_endpoint
for the SDK clients (async client, provider, snapshot SandboxClient) while the
proxy-config PATCH keeps using the root.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit e826864dce0e56cda7decbc48254b1e13eef07e2)

Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
2026-07-17 16:22:38 -04:00
6d34c42578
feat: inject extra JSON fields into sandbox create via env var (#1758)
Add SANDBOX_CREATE_EXTRA_JSON so operators can merge extra fields (e.g.
{"_internal_runtime":"v2"}) into the LangSmith sandbox-create request body.
The SDK's create_sandbox builds a fixed payload with no passthrough, so we
wrap the HTTP client's post to inject the fields on the POST /boxes request
only. Malformed JSON fails at startup validation.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 2238303306493ae6fcd0c2d4ab4236adf283a896)

Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
2026-07-17 16:22:38 -04:00
bfb36be948
feat: add Linear issue search tool (#1748)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 79df6b2ff283afdbd36660be888c130ea964fe3f)

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-07-17 16:22:38 -04:00
ead210927d
fix: defensive copy in get_reviewer_agent and get_chat_agent [closes #1584] (#1742)
* fix: defensive copy in get_reviewer_agent and get_chat_agent [closes #1584]

Factory functions were mutating the caller's RunnableConfig in-place via
config['recursion_limit'] = DEFAULT_RECURSION_LIMIT. Add copy.deepcopy(config)
at the top of each factory and switch the recursion_limit write to setdefault
so a caller-supplied ceiling is respected.

* fix: preserve runtime config object identities

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: Fleet Agent <fleet-agent@langchain.dev>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit c34e04f44da7ec7638cdafc71e0bb40e77070efd)

Co-authored-by: Jacob Albert <122248719+jacobalbert3@users.noreply.github.com>
2026-07-17 16:22:38 -04:00
ae1f883b4c
refactor: move tests into tests/<domain>/ layout
Applies the plan's C5 step: git mv every test per the domain-reorg
move-map (movemap-m50.txt) into tests/{agent,analyzer,auth,dashboard,
github,middleware,models,reviewer,sandbox,slack,tools,webhooks}/, plus
the 13 fork-only placements from the scoping report §2c (Atlassian
webhook tests -> tests/webhooks/, test_atlassian_connect.py and
test_auth_error_leak.py -> tests/auth/, jira/confluence util tests ->
tests/tools/, test_repo_binding_isolation.py -> tests/sandbox/,
bot-identity/autofix tests -> tests/github/).

Path-only move: the only content edits are parents[1] -> parents[2]
fixes in test_e2b_integration.py and test_daytona_integration.py,
required because their __file__-relative ROOT path gained one more
directory level in the move.

Monkeypatch retargets for these files were already completed in C4;
none remained outstanding here.
2026-07-17 14:42:45 -04:00
b3fc62da80
refactor: split webapp.py into api/ + per-source webhook routes
Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved
decisions 1-2): split the 2,590-line agent/webapp.py monolith into
agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py
(composition), agent/api/health.py (/health + /webhooks/run-complete), and
per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian
Connect lifecycle + descriptor routes (/connect/*) fold into
confluence_routes.py; webapp.py becomes the upstream-shaped compatibility
shim (from .api.app import app). langgraph.json http.app stays
agent.webapp:app via the shim.

Fork content, upstream layout: linear/slack route files verified
content-identical to upstream 8356eb34 and taken verbatim; github_routes is
upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are
fork-only, transformed to the same common.X / service.X module-attribute
style. All signature verification (GitHub HMAC, Slack, Linear
timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP
allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo
binding, _is_repo_auto_review_enabled gates, and public-repo org gate move
unchanged.

Handlers rewired from webapp.X to common.X; test monkeypatch sites across
26 files + conftest.py + e2e/harness.py retargeted to
webhook_common/handler/route modules per upstream's pattern. Residual
agent.webapp importers: only the shim, langgraph.json http.app, Makefile
uvicorn target, and docs (doc-path updates land in C7).

Gates: ruff check + format, pytest --co, full unit (1637 passed), full
Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
62d9945df4
refactor: consolidate reviewer modules into agent/review/
Part of the domain-reorg adoption (build plan step C2): fork content,
upstream layout. Nine 1:1 module moves (reviewer_diff/eval_store/
findings/groups/publish/reconcile/trace_context + review_style_
collector/guidance) into agent/review/, with internal relative
imports re-wired to the new package depth. agent/review/__init__.py
mirrors upstream's thin re-export shim (one of the 21 verified "A"
structural adds).

Rewrote the 38 grep hits across importer files (agent/{analyzer,
ci_autofix,reviewer,webapp}.py, agent/dashboard/*, agent/middleware/
settle_review_check.py, agent/tools/*, agent/utils/github_feedback.py,
agent/webhooks/github.py, evals/reviewer/*, and the reviewer test
suite) to point at agent.review.*; 4 of the 38 hits were name
collisions (list_reviewer_findings, reviewer_outcomes,
_reviewer_thread_id, reviewer_thread_id — not the moved modules) and
were left untouched. tests/test_github_checks.py's module-alias
import (`from agent import reviewer_publish`) follows upstream's own
`from agent.review import publish as reviewer_publish` pattern so
downstream `reviewer_publish.*` call sites needed no changes.
agent/reviewer.py and agent/webapp.py stay in place per the hard
rule (fork content, import-only rewire) and are not part of this
package.

Gates: ruff check + ruff format --check, pytest --co -q (1637
collected), full unit suite (1637 passed), and the reviewer/findings
suite in isolation (pytest -k "review or finding", 421 passed).
2026-07-17 13:52:03 -04:00
Adam Moussa
c80bd0f92f
fix: separate review access from automatic reviews (upstream #1720) (#198)
* fix: separate review access from automatic reviews (#1720)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 92d631704e9c38f6aeedf70baede85a7c638567e)

* fix: rename fork-only auto-fix gates to _is_repo_auto_review_enabled

The cherry-pick of upstream 92d63170 renamed _is_repo_enabled_for_review
to _is_repo_auto_review_enabled, but the fork's CI auto-fix, auto-fix
toggle command, and auto-fix review-feedback gates (not present upstream)
still referenced the old name and would raise NameError at request time.
Rename them in place — auto-fix surfaces stay gated on the dashboard
auto-review opt-in list, preserving current fork behavior.

---------

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-07-17 12:20:36 -04:00
Adam Moussa
032d3889e4
fix: align reviewer eval with published findings (upstream #1713) (#201)
* fix: align reviewer eval with published findings (#1713)

* fix: make reviewer eval reflect published findings

Serialize and deduplicate finding persistence, align review calibration around the final six-finding publication, and make judge matching order-independent and auditable.

* fix: honor reviewer eval limits

Forward configured caps into publication snapshots and keep recall-at-cap bounded for diagnostic all-findings runs.

(cherry picked from commit 71e3b8183882bcc42e318f3f220c291617ebcb67)
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* Empty commit to trigger CI

---------

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-07-17 12:10:54 -04:00
Adam Moussa
cfb5624663
feat(open-swe): add E2B sandbox provider (port of upstream 48217b68, #1489) (#199)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
Additive sandbox provider adapted to this fork's synchronous create_sandbox
factory: agent/integrations/e2b.py registered as a lazy-import entry in
SANDBOX_FACTORIES, so the e2b/langchain-e2b imports only load when
SANDBOX_TYPE=e2b (dark-safe on unset env; unset E2B_API_KEY raises a clean
ValueError only when the provider is selected). Supports reconnect-by-id and
optional E2B_TEMPLATE. Non-langsmith providers skip the GitHub proxy step,
so the GitHub-App token flow is untouched.

Upstream: langchain-ai/open-swe 48217b68 (#1489), re-implemented against the
fork's sync sandbox lifecycle rather than cherry-picked (upstream ships on
the deferred async sandbox.py base).

Docs: provider tables/lists in README, CUSTOMIZATION, INSTALLATION; the
CUSTOMIZATION registration example now shows the lazy-tuple form the code
actually uses. uv.lock refreshed (adds e2b, langchain-e2b, dockerfile-parse).
2026-07-16 18:23:42 -04:00
Adam Moussa
22c517a54f
fix: stale admin model defaults after model upgrades (#1709) (#200)
* fix: migrate stale admin model defaults

Normalize retired model IDs before validation and in settings responses so full admin updates remain saveable after model upgrades.

* fix: restrict retired model migration

Only migrate explicitly retired model IDs so malformed provider model names and efforts continue to fail validation.

(cherry picked from commit 62e0ca2d4898ebc3c2ae8887030a364779d907cb)

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-07-16 18:03:44 -04:00
Adam Moussa
7b9eff62e9
fix: enforce terse Slack tool messages (#1717) (#197)
(cherry picked from commit 092abafa4cb955c3823f727d35d7bad94e1147ab)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-07-16 17:51:17 -04:00
Adam Moussa
a534a2e247
chore: let admins interrupt runaway agent runs (#195)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* fix: let admins interrupt runaway agents (#1730)

Add a workspace-wide admin control that cancels every active run while preserving thread history.

(cherry picked from commit 09eaf94c3e612db969daa000b909805963fe1de9)

* chore: reconcile triage ledger

Signed-off-by: Adam Moussa <adam@seahavenind.com>

* fix: change ui import from refactor

Signed-off-by: Adam Moussa <adam@seahavenind.com>

---------

Signed-off-by: Adam Moussa <adam@seahavenind.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-07-16 16:05:03 -04:00
Adam Moussa
0f6eaaea78
chore: adopt five clean upstream cherry-picks (#194)
* chore: include ripgrep in sandbox image (#1728)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 129ddcf9a2fe8b3710535eb10098f80c98b47690)

* feat: include Cargo in sandbox image (#1729)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 1ea03a4330cc9faac65157f3fb94842d074e0231)

* fix: prefer LangSmith tools for trace links (#1751)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit c69459adcafad4a59028232514cecb3bdaf0dfc0)

* fix: add trace link to error banner (#1750)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 5cb2e2bb3582d69241b386bb0852c6f6b40b2dbb)

* fix: link issue PRs and prompt repo conventions (#1704)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 22e024cb1ca080e233eb2cc164767446173a7870)

* chore: reconcile triage ledger

---------

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: Palash Shah <35114859+Palashio@users.noreply.github.com>
2026-07-16 19:28:35 +00:00
Adam Moussa
1f80500cf3
feat: Jira + Confluence integration (tools + triggers) (#182)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat(open-swe): add Jira tool plane (Phase 1)

Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear
tools:

- utils/jira.py: service-account REST client (Basic auth) with get/
  create/update issue, comments, list projects, trace comment; issue and
  comment bodies normalized to markdown.
- utils/adf.py: minimal ADF <-> markdown conversion (read paths convert
  Jira ADF to markdown; agent comments convert prose to ADF).
- tools/jira_{comment,get_issue,get_issue_comments,create_issue,
  update_issue,list_projects}.py wired into the tool registry and the
  main agent tool list.
- tests/test_jira_utils.py: ADF conversion + mocked-transport util tests.

Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env
returns a clean error, so this is safe to land dark. Trigger plane,
prompt guidance, and config plumbing follow in Phase 2.

* feat(open-swe): add Confluence tool plane (Phase 3)

Curated Confluence Cloud REST toolset for the agent, mirroring the Jira
tools:

- utils/confluence.py: service-account REST client (Basic auth) with
  get/create/update page, add comment, CQL search. Page bodies are XHTML
  storage format (not ADF), with minimal storage<->text converters;
  update_page reads the current version and bumps it, as Confluence
  requires.
- tools/confluence_{get_page,create_page,update_page,comment,search}.py
  registered in the tool registry.
- tests/test_confluence_utils.py: converter + mocked-transport tests
  including the version-bump path.

Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN;
unset env returns a clean error. Activation in the agent tool list lands
with the Phase 2 server.py wiring.

* feat(open-swe): add Jira trigger plane (Phase 2)

Make an @openswe comment on a Jira issue spawn an agent run, mirroring
the Linear trigger plane:

- webhooks/jira.py: process_jira_issue clones process_linear_issue —
  deterministic thread id, full-issue fetch, actor accountId->email
  attribution feeding resolve_login_from_email_async (PRs open as the
  human), multimodal image handling, source="jira" + jira_issue config.
- webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time
  X-Automation-Webhook-Token check, fails closed), repo-resolution
  cascade, get_repo_config_from_jira_mapping.
- utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder
  entry — real project->repo mappings still needed).
- utils/jira.py: get_user_email (accountId -> email) for attribution.
- completion.py: source=="jira" failure-reply branch.
- prompt.py: Jira-triggered notify guidance + Refs:/branch key from
  {jira_project_key}-{jira_issue_number}.
- server.py: read jira_issue config + pass jira key to the system
  prompt; also activates the Phase 3 Confluence tools in the agent list.

Jira Automation lacks native webhook HMAC signing, so trust is a shared
secret header (decision D2); replay protection is weaker than Linear's
HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the
outstanding gate/hardening before push.

* fix(open-swe): harden Jira webhook trust (sh-security-review)

Resolves findings from the Phase 2 security review (detector fan-out +
proof-or-kill verifier). The unsigned Jira Automation webhook body was
trusted for identity, comment content, repo routing, and issue
existence; a JIRA_WEBHOOK_SECRET holder could forge those fields.

- Corroborate against the real Jira record: the webhook body is now only
  a pointer (issue_key + required comment_id). The triggering comment's
  author and text are re-fetched server-side via get_comment/fetch_jira_
  comment, and identity, the @openswe check, prompt text, and project
  key are derived from that authoritative record — never payload author/
  body fields. An uncorroborated comment is rejected. (closes the
  account-id impersonation, unsigned-body prompt injection, and
  fabricated-issue findings)
- Validate issue_key against the Jira key format and percent-encode all
  untrusted path segments (_seg) so a crafted key can't traverse to a
  different Jira REST endpoint or inject query params. (closes the path-
  traversal / query-injection findings)
- Route source=="jira" through the bot-token-default / author_prs_as_
  user opt-in path in resolve_github_token, matching Linear, instead of
  unconditionally resolving a per-user OAuth token from a payload email.
- Gate attribution on an active user mapping (is_login_mapped) so a
  pending/unconfirmed mapping can't drive PR authorship.

Adds regression tests: server-corroboration wins over payload, malformed
issue_key rejected, uncorroborated comment rejected, path-segment
encoding, project-key derivation, active-mapping gate.

Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/
REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body +
timestamp on the Automation payload to close the residual replay gap.

* harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist

Folds the two deployment-hardening items from the Phase 2 security review
into code (all opt-in / default-off, so existing and upstream deployments
are unaffected):

- JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must
  carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by
  JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by
  verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear
  HMAC+freshness model). Closes the static-token model's replay/forgery
  gap when enabled.
- JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's
  direct client IP (verify_jira_source_ip). Documented as direct-peer
  only; behind a proxy/LB, allowlist Atlassian's ranges at that layer.
- REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail
  CLOSED instead of the back-compat allow-all, plus a startup fail-open
  warning. Applies to all channels for consistency.

Documents all new vars (and a Jira section) in .env.example. Adds tests
for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and
the fail-closed allowlist.

* feat(open-swe): Confluence Atlassian Connect trigger (Phase 4)

Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian
Connect app. Designed and adversarially verified with the ultracode
workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus
skeptics on the implemented crypto).

- utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's
  official test vector), PyJWT HS256 webhook verifier with alg-pinning,
  issuer binding, and qsh-verified-last ordering; RS256 signed-install
  lifecycle verifier against Atlassian's published keys; installation
  store keyed by clientKey with the sharedSecret encrypted at rest
  (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned).
- webhooks/confluence.py: install/uninstall lifecycle + comment handler.
  The JWT-signed webhook body is only a pointer; the comment's real
  author/text/container are re-fetched server-side via the Basic-auth
  service account (Phase-2 corroboration lesson), with active-only login
  attribution and the repo allowlist.
- utils/confluence.py: get_comment / get_user_email (path-encoded).
- webapp.py: GET /connect/atlassian-connect.json (served dynamically),
  POST /connect/{installed,uninstalled,webhook/comment-created}, the
  space->repo resolver, thread-id, and fetch helpers.
- completion.py: source=="confluence" failure-reply branch.

Security: the sh-security-review verify pass confirmed one HIGH — the
symmetric signed-install=false first-install was trust-on-first-use gated
only by the public Confluence hostname (webhook-auth bypass). Fixed by
switching to signed-install=true + RS256 verification of lifecycle
callbacks, which cryptographically authenticates the first install. All
other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall
DoS/corroboration/injection) were defeated; residuals are deployment
config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover
bodies; comment-trigger prompt injection, shared with all sources).

New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN,
CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets
require the durable Postgres LangGraph store in prod.

Outstanding before push: /sh-security-review on the real diff and the
GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory.

* docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs

- prompt.py: Confluence-triggered runs notify via confluence_comment on
  the triggering page; add Confluence to the shared-base source list.
- CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian
  triggers (Jira Automation shared-secret webhook; Confluence Connect app
  with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the
  server-side corroboration + encrypted install store.

Phase 5 also verified the trigger surface end-to-end against a running
uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed
without valid auth) and recorded the integration in project memory.

* fix(open-swe): resolve /sh-security-review findings on the Atlassian surface

Formal sh-security-review (detector fan-out + verifier) over the Phase-4
Connect surface (esp. the new RS256 signed-install code, unseen by the
earlier adversarial verify) and the Phase-2 opt-in hardening.

CRITICAL — cross-tenant install (origin validation, CWE-346): signed-
install proves the caller is *an* Atlassian tenant, not *ours*, and the
descriptor is served publicly, so any attacker could install the app on
their own Confluence site and drive agent runs against our allowlisted
repos. The baseUrl body field is attacker-controlled and cannot bind the
tenant; only the signature-verified clientKey (JWT iss) can. Added a
MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in
process_install after signature+iss verification.

HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment
ids are per-instance, so generate_thread_id_from_confluence_comment now
salts the hash with the verified clientKey (plumbed from the webhook JWT
iss) to prevent thread hijack across tenants.

HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page
interpolated page_id into the REST path unencoded (update_page on a
mutating PUT with no params= backstop). Now _seg()-encoded, matching the
rest of the module.

MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no
bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID
guard mirroring the Linear botActor / Jira comment_author_is_bot checks.

LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as
opt-in defense-in-depth (the clientKey allowlist is the real gate).

Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp,
kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption,
constant-time comparisons, and the Phase-2 hardening. New regression
tests for each fix; full suite green (1602).

* harden(open-swe): GPT-4.1 cross-family review follow-ups

Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no
critical/high issues and confirmed the auth boundary is fail-closed and
correct. Two low-cost defense-in-depth items applied:

- Validate the signed-install JWT 'kid' against a strict charset before
  the public-key fetch, so a malformed kid fails fast with no network
  call (on top of the existing fixed host + percent-encoding).
- Make JWT nbf verification explicit (verify_nbf) on both the RS256
  lifecycle and HS256 webhook decodes.

Other suggestions triaged as already-handled (aud cross-app replay is
blocked by the per-tenant iss->secret lookup; documented static-token/IP/
baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra
(Fernet rotation via MultiFernet; rate limiting at the gateway).

* docs(open-swe): document Jira + Confluence in installation & customization guides

- INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret,
  service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect
  app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note,
  CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars +
  REQUIRE_REPO_ALLOWLIST.
- CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction
  note covers all four sources.
- AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth).
- README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
seahaven-openswe[bot]
2fb1122630
feat(open-swe): open Linear-triggered PRs as the triggering user (#179)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: open Linear-triggered PRs as the triggering user

Add 'linear' to the set of sources that carry a mapped GitHub login
(Slack, Linear, dashboard, schedule), so Linear-triggered runs resolve
the per-user OAuth token when the author_prs_as_user profile flag is
enabled. Previously only Slack and dashboard runs could open PRs as the
user; Linear runs always used the bot token.

The Linear webhook now resolves the GitHub login from the Linear email
via the same user-mapping store Slack uses, and passes it through the
run configurable and thread owner metadata.

Refs: 5003c953 (upstream #1683)

* fix(open-swe): restrict Linear token attribution to comment author only

Split actor_email (comment_author only, feeds github_login for token
attribution) from user_email (full fallback chain, for display/model).
This ensures a PR is never opened as a non-actor (creator/assignee).

Add test_resolve_github_token_linear_defaults_to_bot to lock the
security-critical default: Linear + mapped login + no opt-in → bot.

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-13 15:17:00 -04:00
seahaven-openswe[bot]
0651de2ebf
feat: surface attributed PR creation failures (#180)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: surface attributed PR creation failures

Port upstream #1659: adds PullRequestCreationGuardMiddleware that
blocks shell fallbacks (gh pr create, gh api /pulls, curl) when
open_pull_request fails, keeping failures visible. Also adds preflight
branch/repo visibility checks in open_pull_request with structured
failure payloads, and updates the prompt to forbid PR creation
fallbacks.

Refs: #134

* fix: fall back to core GitHub App scope when optional grants missing (#1701)

* fix: fall back to core GitHub App scope when optional grants missing

Proxy-token minting requested workflows:write and actions:read in the
permission set used for every sandbox. GitHub 422s a token request that
asks for a permission the installation hasn't granted, so any install
without workflows:write failed to mint a token and every run died in
before-agent setup with "GitHub App installation token is unavailable".

_resolve_proxy_token now walks a permission ladder (full -> +workflows ->
core) and returns the first scope that mints, recording the granted scope
so hourly proxy refreshes stay consistent. A missing optional grant now
degrades to the install-time core scope instead of failing the run;
workflow-file HITL pushes still require workflows:write and fail at push
time when it is absent.

* refactor: flatten proxy-token ladder loop with continue

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit f53caff1aa24a7b29d851b267aa3bdfe62c1e935)

Sea Haven fork deviation: upstream #1701 folds workflows:write into the
standing BASE/RUNTIME scope. This fork deliberately keeps workflows:write
OUT of the standing permission ladder (RUNTIME = core + actions:read;
LADDER = (RUNTIME, CORE)) so the sandbox proxy token cannot push
.github/workflows/* during normal operation. workflows:write is minted only
transiently by WorkflowPushGuardMiddleware for an approved HITL push and
dropped on restore, preserving token scope as a backstop for the workflow-
push approval control. Security-reviewed (agentic fan-out + GPT-4.1 cross
review); the standing-scope-carries-workflows:write bypass was blocked.

* fix(open-swe): harden proxy-token restore and mint error handling

Two low-severity follow-ups from the security review of the #1701 port.

Restore the recorded baseline scope after a workflow-push elevation instead
of a hardcoded RUNTIME. An install granted workflows:write but not actions:read
resolves its standing token to core; hardcoding RUNTIME on restore requested the
ungranted actions:read, 422'd, and fired a false "SECURITY: failed to downscope"
error on every approved workflow push before the core fallback recovered. The
guard now captures the run's recorded scope before elevating (via the new
get_recorded_proxy_permissions) and restores exactly that, falling back to the
guaranteed core scope only when the baseline restore fails.

Classify installation-token mint failures. get_github_app_installation_token_
with_expiry now treats HTTP 422 (a permission the installation hasn't granted)
as the ladder's expected descend signal and keeps it at debug, while a non-422
failure (network/5xx/timeout) is surfaced at WARNING even when errors are
otherwise suppressed — so a transient blip no longer silently downscopes a whole
run under a debug-only trace. The reduced-scope warning no longer asserts a
missing grant as the sole cause.

* chore(triage): mark upstream #1701 landed on this branch

Ported via PR #181 as Option A (workflows:write kept out of the standing
proxy-token scope). Regenerated triage.md from triage.jsonl.

* fix: restructure PR creation to POST-first with diagnose-on-failure

Move preflight checks from an authoritative gate (before POST) to a
diagnostic run after POST failure. This avoids false-positive failures
when a just-pushed head branch is momentarily invisible to GitHub ref
endpoints, and eliminates 2-3 extra serial API round-trips on the happy
path.

Also drop unused _PR_CREATED_FALSE indirection and add a docstring to
pr_creation_guard acknowledging the fail-open detection design.

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-13 14:45:19 -04:00
Adam Moussa
e9da186b5f
fix(open-swe): port core GitHub-App scope fallback (#1701), workflows:write kept out of standing scope (#181)
* fix: fall back to core GitHub App scope when optional grants missing (#1701)

* fix: fall back to core GitHub App scope when optional grants missing

Proxy-token minting requested workflows:write and actions:read in the
permission set used for every sandbox. GitHub 422s a token request that
asks for a permission the installation hasn't granted, so any install
without workflows:write failed to mint a token and every run died in
before-agent setup with "GitHub App installation token is unavailable".

_resolve_proxy_token now walks a permission ladder (full -> +workflows ->
core) and returns the first scope that mints, recording the granted scope
so hourly proxy refreshes stay consistent. A missing optional grant now
degrades to the install-time core scope instead of failing the run;
workflow-file HITL pushes still require workflows:write and fail at push
time when it is absent.

* refactor: flatten proxy-token ladder loop with continue

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit f53caff1aa24a7b29d851b267aa3bdfe62c1e935)

Sea Haven fork deviation: upstream #1701 folds workflows:write into the
standing BASE/RUNTIME scope. This fork deliberately keeps workflows:write
OUT of the standing permission ladder (RUNTIME = core + actions:read;
LADDER = (RUNTIME, CORE)) so the sandbox proxy token cannot push
.github/workflows/* during normal operation. workflows:write is minted only
transiently by WorkflowPushGuardMiddleware for an approved HITL push and
dropped on restore, preserving token scope as a backstop for the workflow-
push approval control. Security-reviewed (agentic fan-out + GPT-4.1 cross
review); the standing-scope-carries-workflows:write bypass was blocked.

* fix(open-swe): harden proxy-token restore and mint error handling

Two low-severity follow-ups from the security review of the #1701 port.

Restore the recorded baseline scope after a workflow-push elevation instead
of a hardcoded RUNTIME. An install granted workflows:write but not actions:read
resolves its standing token to core; hardcoding RUNTIME on restore requested the
ungranted actions:read, 422'd, and fired a false "SECURITY: failed to downscope"
error on every approved workflow push before the core fallback recovered. The
guard now captures the run's recorded scope before elevating (via the new
get_recorded_proxy_permissions) and restores exactly that, falling back to the
guaranteed core scope only when the baseline restore fails.

Classify installation-token mint failures. get_github_app_installation_token_
with_expiry now treats HTTP 422 (a permission the installation hasn't granted)
as the ladder's expected descend signal and keeps it at debug, while a non-422
failure (network/5xx/timeout) is surfaced at WARNING even when errors are
otherwise suppressed — so a transient blip no longer silently downscopes a whole
run under a debug-only trace. The reduced-scope warning no longer asserts a
missing grant as the sole cause.

* chore(triage): mark upstream #1701 landed on this branch

Ported via PR #181 as Option A (workflows:write kept out of the standing
proxy-token scope). Regenerated triage.md from triage.jsonl.

---------

Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
2026-07-13 14:24:37 -04:00
seahaven-openswe[bot]
00a99c86c4
chore: port question-answering prompt clarification (#1680) (#178)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
Port upstream 304032fa: clarify that information-only answers should
check out relevant repos first for context, answer fully inline, and
only post a concise summary to Slack threads.

The fork already carried the shared-base Slack guidance; this adds the
missing TASK_EXECUTION_SECTION update and its test coverage.

Refs: #146

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-13 13:16:26 -04:00
Adam Moussa
23cf6d2d4d
fix(open-swe): make workflow-push-guard blocks tell the agent how to recover (#173)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
The WorkflowPushGuardMiddleware fails closed and blocks a git push whenever
it cannot parse the command into a plain, inspectable form (chained shell
operators, obfuscation, unrecognized refspecs). Most block reasons named only
the symptom ("chained shell commands"), so the agent kept retrying other
chained variants (cd &&, pushd &&) instead of dropping the chaining.

Append a single actionable remedy to every block reason, pointing at a plain
`git push origin <branch>` or `git -C <dir> push`, so a blocked run recovers
on the next attempt instead of looping.
2026-07-10 15:35:06 -04:00
Adam Moussa
5350b63c85
feat(open-swe): re-add Fable 5 behind an admin toggle (Bedrock) (#172)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat(models): re-add Fable 5 with admin disable toggle (port of upstream #1677)

* refactor(models): convert re-added Fable 5 to Bedrock model IDs

* fix(open-swe): correct Fable copy to describe provider data sharing, not ZDR

The ported admin toggle description and code comments described Fable 5 as
incompatible with Zero Data Retention. That is backwards: Fable 5 requires
the account to opt into Bedrock provider_data_share — prompts/completions are
retained and shared with Anthropic (up to 30 days, incl. human review). The
old UI copy would lead an admin to believe the opposite of what enabling the
toggle does. Reword the toggle description and the gate_fable_model /
team_settings comments accordingly. Still off by default. Refs #171.
2026-07-10 14:05:20 -04:00
Adam Moussa
bc59c05642
feat(dashboard): surface thread sandbox ID in sidebar (#169)
* feat(dashboard): surface thread sandbox ID in sidebar

Port of upstream #1689 (feb7ac98): expose the thread's sandbox_id on the
dashboard thread summary and surface it in the sidebar via a "Copy sandbox
ID" action.

- thread_api: add sandboxId to the thread summary, hiding the
  "__creating__" in-flight sentinel.
- queries/types: thread the sandboxId field through AgentThread and the
  optimistic thread.
- AgentsSidebar: replace the hover-only resolve/delete buttons and the
  right-click context menu with one touch-friendly kebab (⋮) menu that
  works on both pointer and touch, and add the Copy sandbox ID item.
- vite: disable the PWA service worker in dev (it precaches assets and
  defeats HMR).
- tests: unit test for the summary field + e2e spec in the real-backend
  harness.

Closes #140. Part of #134.

* chore(triage): mark feb7ac98 (#1689) landed

Ported to dev via feat/thread-sandbox-id-sidebar (#140).
2026-07-10 13:17:03 -04:00
seahaven-openswe[bot]
e3c7c03c92
feat: port model-fallback resilience from upstream (#1694, #1695) (#161)
* feat: port plan-review & workflow-approval UX (#135)

Port six upstream commits onto dev:

- c03a6be7 (already ported): keep plan guidance high-level
- 546042a4: add workflow approval UI with diff preview, approval URLs,
  web review links, and polling for approval status during active runs
- 216cf181: remove workflow token elevation; approved pushes pass
  through directly without proxy token rewriting
- 3dbc0282: preserve plan redirects after login by accepting relative
  same-origin redirect_to values and rejecting blocked paths
- bb104d93: submit plan comments with cmd+enter
- 90cb6caa: terse Slack replies, shared content via save_plan outside
  plan mode (PLAN_STATUS_SHARED), reject shared-content mutations

Refs: #135

* feat: port durable dispatch hardening and startup latency improvements

Port five upstream PRs onto dev:

- #1621 / #1658: durable dispatch with loopback webhook defense,
  create_durable_run helper, _config_with_prepare_run_id, degradation
  to None for relative/loopback completion webhook URLs
- #1696: run-level completion webhook deduplication (replace
  claim-then-post with post-then-flag per run_id), DeferredErrorModel
  for graph-factory resilience, ToolRetryMiddleware for task subagents,
  TimeoutWrapupMiddleware for all three graphs
- #1697: lazy-load __init__.py for agent.middleware, agent.tools,
  agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search,
  agent.webapp in request_pr_review, deepagents in sandbox.py); add
  ttl_cache.py with stale-while-revalidate for tool loaders

Refs: #137

* fix: restore login page render and clear CI lint/format

The plan-review port removed the authRedirectUrl import from login.tsx
but left its call site, crashing the login page at runtime (blank page,
no 'Sign in to open-swe'). Pass the relative path straight to loginUrl,
matching the plan route and the backend relative-redirect handling.

Also drop an unused os import in the guard test and reformat
workflow_push_guard.py to satisfy ruff.

* feat: port model-fallback resilience from upstream (#1694, #1695)

- Add httpx.TransportError to the transient-exception set so
  incomplete chunked reads on streamed responses trigger a
  fallback instead of cancelling the run (#1694 / c9f6dd86).
- Rewrite fallback to alternate primary/fallback with exponential
  backoff instead of a single failover, so the agent survives
  multi-minute gateway outages spanning both providers (#1695 /
  c9a9a7cd).
- Default backoff schedule (0, 5, 15, 30, 45) reaches past the
  gateway's ~30s recovery window; jittered ±25%.
- On exhaustion, surface a terminal AIMessage explaining the
  outage instead of crashing — progress is checkpointed so the
  user can retrigger to continue.
- Preserve fork conventions: Bedrock ClientError retryability
  check, sync wrap_model_call (using time.sleep instead of
  asyncio.sleep), and existing access-error surfacing for
  Anthropic, OpenAI, and Bedrock (botocore) provider errors.
- Update triage ledger (c9f6dd86, c9a9a7cd → landed) and
  re-render triage.md.

Refs: #139

* fix: align workflow-push-guard tests with dev's transient-elevation impl

The dev merge auto-combined dev's elevation tests with the stale passthrough
tests inherited from the durable-dispatch branch; the passthrough tests
contradict dev's restored _run_with_workflow_token impl. Take dev's test file.

* fix: drop dead ttl_cache module; make fallback backoff jitter two-sided

ttl_cache.py was re-introduced via the dev merge but dev/#160 deliberately
removed it as dead code (no agent importer). Remove it to match dev.
Also make _jittered_delay symmetric (±25%) to match its docstring.

* chore(upstream-sync): triage 4 new upstream commits (#1708-#1713)

Synced ledger to upstream/main (71e3b818). New rows all deferred:
- #1708 add GPT-5.6 OpenAI models (FLAG-HUMAN: fork picker is Bedrock/Fireworks-only)
- #1709 stale admin model defaults after upgrades
- #1710 bump langchain-fireworks 1.4.4
- #1713 align reviewer eval with published findings

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 18:28:21 -04:00
seahaven-openswe[bot]
0546085672
feat: port durable dispatch hardening and startup latency improvements (#160)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: port plan-review & workflow-approval UX (#135)

Port six upstream commits onto dev:

- c03a6be7 (already ported): keep plan guidance high-level
- 546042a4: add workflow approval UI with diff preview, approval URLs,
  web review links, and polling for approval status during active runs
- 216cf181: remove workflow token elevation; approved pushes pass
  through directly without proxy token rewriting
- 3dbc0282: preserve plan redirects after login by accepting relative
  same-origin redirect_to values and rejecting blocked paths
- bb104d93: submit plan comments with cmd+enter
- 90cb6caa: terse Slack replies, shared content via save_plan outside
  plan mode (PLAN_STATUS_SHARED), reject shared-content mutations

Refs: #135

* feat: port durable dispatch hardening and startup latency improvements

Port five upstream PRs onto dev:

- #1621 / #1658: durable dispatch with loopback webhook defense,
  create_durable_run helper, _config_with_prepare_run_id, degradation
  to None for relative/loopback completion webhook URLs
- #1696: run-level completion webhook deduplication (replace
  claim-then-post with post-then-flag per run_id), DeferredErrorModel
  for graph-factory resilience, ToolRetryMiddleware for task subagents,
  TimeoutWrapupMiddleware for all three graphs
- #1697: lazy-load __init__.py for agent.middleware, agent.tools,
  agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search,
  agent.webapp in request_pr_review, deepagents in sandbox.py); add
  ttl_cache.py with stale-while-revalidate for tool loaders

Refs: #137

* fix: restore login page render and clear CI lint/format

The plan-review port removed the authRedirectUrl import from login.tsx
but left its call site, crashing the login page at runtime (blank page,
no 'Sign in to open-swe'). Pass the relative path straight to loginUrl,
matching the plan route and the backend relative-redirect handling.

Also drop an unused os import in the guard test and reformat
workflow_push_guard.py to satisfy ruff.

* fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary

* fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache

- Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__
  (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it.
- Reroute E2E model patching to deferred_model.make_model so make_model_or_defer
  (used by all three graph factories) returns the scripted fake instead of
  building a real model with fake credentials.
- Drop unused agent/utils/ttl_cache.py — no agent module imports it.
- Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001).
- Format tests/test_dispatch.py.

* fix: claim-then-post run-level failure dedup; stop permanent suppression

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00
seahaven-openswe[bot]
f87847baa4
feat: port plan-review & workflow-approval UX (#159)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: port plan-review & workflow-approval UX (#135)

Port six upstream commits onto dev:

- c03a6be7 (already ported): keep plan guidance high-level
- 546042a4: add workflow approval UI with diff preview, approval URLs,
  web review links, and polling for approval status during active runs
- 216cf181: remove workflow token elevation; approved pushes pass
  through directly without proxy token rewriting
- 3dbc0282: preserve plan redirects after login by accepting relative
  same-origin redirect_to values and rejecting blocked paths
- bb104d93: submit plan comments with cmd+enter
- 90cb6caa: terse Slack replies, shared content via save_plan outside
  plan mode (PLAN_STATUS_SHARED), reject shared-content mutations

Refs: #135

* fix: restore login page render and clear CI lint/format

The plan-review port removed the authRedirectUrl import from login.tsx
but left its call site, crashing the login page at runtime (blank page,
no 'Sign in to open-swe'). Pass the relative path straight to loginUrl,
matching the plan route and the backend relative-redirect handling.

Also drop an unused os import in the guard test and reformat
workflow_push_guard.py to satisfy ruff.

* fix: carry workflows:write on the standing proxy token

Complete the half-ported upstream 216cf181 cascade. The port dropped
_run_with_workflow_token from the guard but missed the paired github_app
change, so an approved .github/workflows push ran with the base token
(no workflows:write) and GitHub 403'd it.

Add workflows:write to BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS and delete the
now-orphaned WORKFLOW_RUNTIME_PROXY_TOKEN_PERMISSIONS constant; update the
github_app and proxy_auth tests to match. The HITL approval gate in
workflow_push_guard.py is unchanged — this only lets the standing token
push once a human approves.

* fix: restore transient workflow-token elevation (revert standing workflows:write)

The standing GitHub-App proxy token (BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS) is
ALWAYS-ON, so carrying workflows:write on it made the fork's HITL workflow-push
guard the sole control over unapproved workflow pushes. The guard's git-push
parser has gaps (obfuscated-expansion push, `gh api` REST contents PUT,
fully-qualified cross-branch refspecs); with a permanently workflows-scoped
token those gaps become live unapproved-workflow-push exploits (1 critical, 2
high — security review BLOCK on #159).

Restore dev's transient-elevation model:
- Drop workflows:write from BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS; re-add the
  WORKFLOW_RUNTIME_PROXY_TOKEN_PERMISSIONS constant (base + workflows:write).
- Re-introduce _run_with_workflow_token in the guard: it mints the
  workflows-scoped token via refresh_proxy_token around the approved,
  guard-normalized fixed_command, then downscopes to RUNTIME then BASE in a
  finally. Route the approval branch through it.
- Restore the dev token/elevation tests.

The standing token no longer carries workflows:write, so the three parser
bypasses hit GitHub 403 again; an approved push still succeeds because the
elevation grants workflows:write only around the normalized command. Keeps all
of #159's diff-preview / approval-URL / Slack-card guard additions.

* fix: reject protocol-relative path from sanitizeAuthRedirect (open redirect)

sanitizeAuthRedirect returned parsed.pathname+search+hash, which `new URL` can
resolve to a protocol-relative `//host` (e.g. input `/..//evil.com` normalizes
same-origin, passing the origin check, but yields a path starting with `//`).
ClientRedirect / login.tsx feed that path to window.location.replace, so it
navigates cross-origin — an open redirect. Reject any resolved path that is not
a single-leading-slash path (`^/[^/]`), falling back to the default. Adds
coverage for `/..//evil.com`, `/.//evil.com`, and `//evil.com`.

* fix: log SECURITY error when workflow-token downscope fails

The elevate->push->downscope finally block was silent on failure. If both
refresh_proxy_token calls fail, the sandbox retains workflows:write for the
rest of the run with no signal. Log a SECURITY error on the partial and full
downscope-failure paths so the retention is observable.

Addresses the GPT-4.1 cross-family review of the token-scope remediation.

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 16:03:13 -04:00
seahaven-openswe[bot]
c2bc7720cd
feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678)

Ports four upstream commits that add opt-in LLM call routing through the
LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs,
no-agent-attribution, bun toolchain).

- #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings
  toggle, admin UI section, wired into make_model for all graph entrypoints
- #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over
  platform LANGSMITH_API_KEY
- #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host,
  SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware
- #1678 (73b7d1c0): fix OpenAI Responses reasoning replay —
  SanitizeOpenAIResponsesMiddleware, store/include config for encrypted
  reasoning content, reasoning_effort coercion for Chat Completions fallback

Refs #134

* fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity

- Downgrade logger.warning to logger.debug in gateway_overrides for
  not-routed providers and missing API key (Bedrock is the default
  provider in this fork, so these are expected steady states)
- Add Bedrock to the LLMGatewaySection route-toggle description so
  admins know it is not routed through the gateway
- Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with
  server.py and reviewer.py
- Restore the Bedrock region comment in model.py that explains the
  AWS_REGION / AWS_DEFAULT_REGION precedence

Refs #138

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-09 14:44:15 -04:00
Adam Moussa
b65a99f3e9
feat: re-land scoped AGENTS auto-load (#1684) + platform issue reporting tool (#1685) (#129)
* feat: auto-load scoped AGENTS on reads (#1684)

Adapted from upstream langchain-ai/open-swe #1684 to the fork's
direct-import middleware registry and middleware stack ordering.
Adds SubdirAgentsReadMiddleware, which appends applicable ancestor
AGENTS.md instructions to read_file results once per run, so scoped
rules are visible before edits. Wired into get_agent immediately after
ToolErrorMiddleware, matching upstream's relative position.

Note: this changes file-read behavior for every main-agent run. The
reviewer graph uses its own leaner middleware stack and is unaffected.

(cherry picked from commit 7f7af71547be2199cea699284676f8ceefba7691)

* feat: add platform issue reporting tool (#1685)

Adapted from upstream langchain-ai/open-swe #1685 to the fork's
direct-import tool registry. Adds the report_platform_issue tool
(stdlib-only: returns a locally generated UUIDv7 report id, no external
network call) and wires it into get_agent's curated tool list.

Dropped upstream's test_task_retry_wraps_inside_tool_error_middleware
assertion, which references ToolRetryMiddleware that this fork does not
wire into the middleware stack.

(cherry picked from commit 88b62322b44103335773002d0745704cd96e9160)

---------

Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
2026-07-08 18:52:01 -04:00
seahaven-openswe[bot]
a53a96d37b
feat: editable plan mode — owner hand-edits the plan before approval (#1610, #80) (#130)
* feat: editable plan mode — owner hand-edits the plan before approval (#1610, #80)

Adds a PUT /dashboard/api/plan/{thread_id} endpoint and PlanReview UI edit
mode so the thread owner can refine the published plan markdown by hand. The
edited markdown is re-published as "ready" (preserving reviewer comments),
mirrored into the sandbox plan.md, and handed to the agent as the source of
truth on approve.

Approve now reads the published plan content strictly (raise_on_error=True)
so a transient store failure aborts instead of silently dropping an edited
plan, matching the comment-read contract.

The banner-overlap fix (collapsed git-panel clearing the "Review plan" link)
was already ported in #128; this picks up the remaining edit-mode pieces.

Refs #80

* fix(plan): make approve_plan idempotent, dispatch before persisting, fix comment count

SH-128-03: approve_plan set status APPROVED before dispatching the follow-up
run and had no already-approved guard, so a failed dispatch left the plan stuck
approved-but-undispatched and a double-submit double-dispatched + double-posted
the Slack notice. And the Slack notice counted len(comments) including empty
comments _format_comments filters out.

- Return 409 when the plan is already approved (idempotent double-click/retry).
- Dispatch the implementation run BEFORE persisting APPROVED so a dispatch
  failure leaves the plan re-approvable. _dispatch_followup passes plan_mode
  explicitly, so the run is unaffected by the reorder.
- Count only non-empty comments in the Slack approval notice.

Fixed here (not on #128) because #128's approve_plan is rewritten on this
branch; #129 inherits it. Adds tests for the 409, the filtered count, and the
dispatch-before-status ordering.

---------

Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-08 18:44:56 -04:00
Adam Moussa
b3b0274403
feat: Re-land deferred upstream features on modular webhooks (#80) (#128)
* fix(webhooks): fall back to vision model for Slack/Linear image threads

Re-land upstream #1626 onto the modular webhook structure. When a
Slack mention or Linear issue carries images but the resolved model is
text-only, fall back to a vision-capable model instead of dropping the
images. Re-points default_vision_model_pair at the fork's image-capable
models (Opus 4.8 default, else any supports_images model) rather than
upstream's openai:/anthropic: provider filter.

Refs #80, upstream #1626

* fix(slack): persist trace_message_ts so web-handoff updates the trace reply

Re-land upstream #1630 onto the modular structure. The first-mention
store_slack_run_mapping call did not pass trace_message_ts, so it was
never persisted (nothing to preserve from on first mention) and
_notify_slack_web_handoff always skipped the trace-reply update on web
handoff. Pass it through and cover it with a test.

Refs #80, upstream #1630

* feat(slack): include channel context in Slack prompts

Re-land upstream #1633 onto the modular structure. Fetch cached Slack
channel metadata once per event (_get_slack_channel_context) and thread
it through the docs-plz gate, repo resolution, and process_slack_mention
so prompts carry the channel name and a clearly-marked untrusted
channel description. Avoids duplicate conversations.info calls.

Refs #80, upstream #1633

* feat(tools): add slack_start_new_thread breakout tool

Re-land upstream #1638 onto the modular structure. Adds the
slack_start_new_thread tool (posts a top-level Slack message and
dispatches a fresh agent run for a broken-out task via the durable
dispatch_agent_run contract), wires it into the agent tool list and
tools/__init__, adds prompt guidance, and excludes it from plan mode so
it can't bypass the approval flow. Tool imports only live modules.

Refs #80, upstream #1638

* feat(plan): notify Slack on plan approval

Re-land upstream #1632 onto the modular structure. When a plan is
approved via the dashboard approve endpoint, post a thread reply to the
originating Slack thread noting the comment count and approver, after
the follow-up run is dispatched. Slack post failures never break
approval. Adapted to the fork's approve_plan (no plan_markdown read).

Refs #80, upstream #1632

* feat(plan): publish plans from sandbox files

Re-land upstream #1635 onto the modular structure, completing the
partially-ported change so dev is internally consistent. save_plan now
takes a plan_file_path, reads the agent-authored Markdown file from
/workspace/plans/ (validating extension/location/UTF-8/size) and
publishes it, instead of taking a plan_markdown string. Removes
write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can
author the plan file, updates enter_plan_mode/reject_plan guidance and
the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev).

Refs #80, upstream #1635

* fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs

INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop
revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/
Linear image URL could 302-redirect the fetch to an internal host / cloud
metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot
is_url_safe check. Route image fetches through the same per-hop resolve+pin+
revalidate loop the http_request tool uses, lifted into url_safety as the shared
request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token
on redirect so it can't be replayed to a redirect target.

SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at
DEBUG; multimodal logged them at INFO on every fetch. Log host-only.

Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened
its reach by no longer dropping images for text-only models. Fixing on the base
branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression
tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00
ea1845d41d
fix(security): fence and neutralize untrusted Slack channel description in agent prompt
Hardens SLACK-PI-001 (sh-security-review). Slack channel topic/purpose is
editable by ordinary channel members and flowed verbatim into the agent LLM
prompt behind only a prose 'untrusted' label — an indirect prompt-injection
vector for an agent with network egress and repo write. Now strip leading
markdown structural tokens per line (so it can't forge the prompt's real
request/section delimiters), cap length, and wrap it in a per-render
unguessable sentinel fence (so injected text can't spoof a closing marker to
escape the data block). Deliberately diverges from upstream #1633.
2026-07-03 16:08:35 -04:00
Johannes du Plessis
e8b6fb7050
fix: surface Slack thread errors (#1627)
* fix: surface Slack thread errors

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't set failure_reply_posted on Slack preprocessing errors

The preprocessing error handler was setting failure_reply_posted=True,
the same idempotency flag handle_run_completion checks to suppress
duplicate run-failure replies. Since preprocessing failures happen
before any run exists but the flag persists on the thread, a subsequent
run failure on the same thread would be silently ignored.

The preprocessing handler already posts its own Slack reply, so the
run-completion idempotency flag should not be set here.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit bb36448b0b)
2026-07-03 15:50:30 -04:00
Ramon Nogueira
3c6077c418
feat: add Slack reaction tool (#1650)
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit ee224d3e91576f771e93df7bad4523a7a7036324)
2026-07-03 15:38:56 -04:00
Ramon Nogueira
b2b49ed944
feat: add Slack breakout thread tool (#1638)
* feat: add Slack breakout thread tool

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: make fake LLM scripts declarative

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: exclude slack_start_new_thread from plan mode

The breakout tool can dispatch a fresh agent run that starts outside the
current plan-mode state, bypassing the approval flow. Add it to
PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating
tools while planning.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 747ce4bbe5)
2026-07-03 15:37:13 -04:00
5dca57146b
fix: persist trace_message_ts on Slack run mapping
Hand-applied residual of upstream #1630 (rest already integrated).
2026-07-03 14:59:31 -04:00
Ramon Nogueira
3a941a693c
feat: include Slack channel context in prompts (#1633)
Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
(cherry picked from commit 27d90ef196)
2026-07-03 14:56:43 -04:00
Adam Moussa
c4905ac01e
fix(models): make Claude family fallback recognize Bedrock ids (#119)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
#1651 added the family-aware `provider_fallback_pair` via `_claude_family_of`,
but the helper only matched `anthropic:claude-*` ids. This fork serves Claude
through Bedrock (`bedrock_converse:us.anthropic.claude-*`), so the family logic
was dead code: a dropped Bedrock Sonnet fell back to the Bedrock Opus that sits
first in the list instead of staying in the Sonnet family.

Teach `_claude_family_of` to parse `bedrock_converse` ids and add regression
coverage for the Sonnet-stays-on-Sonnet case.
2026-07-03 11:51:41 -04:00
Adam Moussa
589cd236c6
chore: cherry-pick clean upstream fixes + cherry-pick runbook (#117)
* fix: make plan view mobile friendly (#1636)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 7ee3e05724)

* fix: return to thread after plan approval (#1637)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit f32e492ab4)

* feat: reviews block agenda, sticky headers, accurate diff scroll (#1653)

Rework the AI-sorted blocks experience on the PR reviews page into a
Google-Docs-style outline: the left sidebar is now a clean number+title
agenda with scroll-spy highlighting of the active block; each block shows
its title + description (sticky) above its diff; and diff rows are pinned to
a uniform height so scroll-to lands precisely via the virtualizer's own
geometry instead of an estimate-driven correction loop.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 0b76afdc955e33805c7623d1502a75a9c7c9c1b7)

* fix: jump + ResizeObserver settle for review scroll-to (#1655)

Replace smooth-scroll plus frame-count correction loops on the PR
reviews page with an instant jump that re-asserts its target via a
ResizeObserver (the real "layout settled" signal). Block/file
navigation and finding/comment centering now land deterministically as
off-screen cards mount, files expand, and annotation cards measure,
instead of racing a smooth-scroll animation against height
reconciliation. Holds bail on user wheel/touch input and after a short
ceiling, and a new navigation cancels the previous hold.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
(cherry picked from commit 7530653bba7774d66a54b8bef0d2bbc25f519942)

* fix: purge expired thread_wakeup crons (#1656)

* fix: purge expired thread_wakeup crons

One-shot wakeup crons set an end_time that stops re-firing but the cron
row is never deleted, so dead rows accumulate (86 in prod). Add a purge
that deletes thread_wakeup crons past their end_time, called
opportunistically before scheduling a new wakeup, plus a one-time
backfill script. Conservative: matches only kind=thread_wakeup with a
past end_time.

* chore: retrigger Open SWE review

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 9e5a1924ef306269322c31342a1831e57831cfee)

* fix: add top padding to sticky review block header (#1660)

* fix: add top padding to sticky review block header

The sticky per-block header on the reviews page had padding below but
none above, so the block number badge sat glued against the top edge
when pinned. Add matching top padding for breathing room.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: use py-2 shorthand for review block header padding

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 23bd4a63fc5ba0fe853babf79ed33feb866cc8b2)

* fix: use global tokens for sidebar filter popover border (#1661)

The filter popover renders via base-ui Menu.Portal into document.body,
outside the .agents-ui container where the --ui-* CSS variables are
scoped. As a result border-[var(--ui-border)] resolved to an undefined
variable and border-color fell back to currentColor, producing a strong
near-black border (separators/hover/labels were similarly off).

Switch the portaled popup styling to the same global shadcn tokens the
theme/settings popover (SidebarUserMenu) already uses (border-border,
bg-border, bg-muted, text-muted-foreground). These are defined at :root
so they resolve inside portals too, and match the settings popover.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 63eb9a08209f683016abf01cdcc548bc5905f158)

* fix: preserve dashboard redirect after login (#1668)

* fix: preserve dashboard redirect after login

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: cover plan login redirect in e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit bc7ce59169b5350da7286164afb83a7b037b528d)

* Disable React StrictMode (#1654)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 6575c327a3ac2b107a6e79a04fa61168d779dbf0)

* docs(upstream-sync): add cherry-pick runbook

Repo-specific runbook for bringing upstream (langchain-ai/open-swe) commits
into the fork: triage-sync discovery, the git cp workflow, the triage ledger,
themed-branch layout, and conflict/regression handling.

---------

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com>
2026-07-03 11:48:40 -04:00
Adam Moussa
b65c3a07db
feat: distill Sea Haven conventions into agent prompt, reviewer, and fork docs (#113)
* Add Dependabot ignore for @types/node semver-major bumps

Prevent Dependabot from proposing wrong-direction @types/node major
bumps (e.g. 24 -> 26). /ui runs on Node 24 on Vercel; a too-new types
major still compiles but describes APIs absent at runtime.

Refs: #110

* feat(agent): seed all-repos custom instructions in default_prompt.md

Distill the universally-applicable Sea Haven authoring conventions into the
team-default Custom Instructions the main agent gets on every repo: secrets/
config placement, keep-docs-in-sync, verify-before-push, re-run-real-gates
after delegating, confirm-a-convention-before-adopting, and house writing
style. Toolchain references are generalized (not tied to a specific stack).

* feat(reviewer): seed Sea Haven review baseline as org-guidelines default

Bake DEFAULT_ORG_REVIEW_GUIDELINES (severity model, secrets, security surface,
tests, naming, deferred-work-needs-an-issue) and default org_guidelines to it
in _default_settings(). The reviewer now applies the Sea Haven baseline on
every repo until a workspace admin overrides it with a non-empty value via
the dashboard. Stack-agnostic and well under the 10k-char cap.

* refactor(prompt): consolidate duplicated COMMIT_PR_SECTION + add fork-sync runbook

COMMIT_PR_SECTION had two overlapping passes with a contradictory PR-title
rule (a fixed 'type: description' form vs the repo-aware detection). Collapse
into one numbered sequence (lint -> commit -> push/PR -> notify), keep the
authoritative repo-aware title rule, and drop the duplicate notify step. All
IMPORTANT directives (force-push ban, workflow-approval, autonomy, 403
handling) are preserved verbatim.

Add a fork-maintenance runbook to CLAUDE.md distilling the durable
upstream-sync methodology (conflict triage, deferred-refactor resolution rule,
the silent re-import/wiring hazards, test-impl-same-side, layered CI).

* refactor(prompt): adopt conventional-commit style

Flip the Sea Haven authoring convention baked into the agent prompt from
imperative/no-prefix to conventional-commit style:
- Commit subjects and the no-gate PR-title default now use
  type(scope): description with the allowed type set (feat, fix, docs,
  style, refactor, perf, test, build, ci, chore, revert, release).
- Branch prefixes expanded to feature/, fix/, hotfix/, chore/, docs/,
  refactor/, release/ (kebab-case description).
- The repo-aware gate detection is preserved: a repo's own title gate
  still wins and may narrow the allowed types/scopes.

Updated test_github_comment_prompts.py to assert the new convention.
2026-07-02 18:29:15 -04:00
seahaven-openswe[bot]
47a9a1ba89
feat: Loosen workflow push approval fingerprint to repo/branch/files identity (#101) 2026-07-01 21:11:31 -04:00