Commit graph

70 commits

Author SHA1 Message Date
open-swe[bot]
d2c62e5fc6
chore: remove slack assistants api feature flag (#1295)
The SLACK_ASSISTANTS_API_ENABLED env flag and its gating function
are removed. set_slack_assistant_status now always proceeds when a
bot token and channel/thread are provided.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-12 00:41:51 +00:00
Johannes du Plessis
5c7c78406c
fix: keep sandbox backend stable across recovery (#1294)w
Use a per-thread proxy so in-flight tools continue through the latest recreated sandbox instead of holding a stale backend reference.
2026-05-11 16:03:38 -07:00
open-swe[bot]
ed7974f62d
feat: remove eyes reaction on Slack invocation (#1289)
Drops the add_slack_reaction call in process_slack_mention; the Slack
assistant status indicator and the agent's first reply still signal
acknowledgement. Linear and GitHub eyes reactions are unchanged.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
2026-05-11 15:07:02 -07:00
open-swe[bot]
85343fab63
feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322] (#1280)
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]

Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* webapp: forward installation-token expiry to reviewer cache writes

The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 22:57:01 +00:00
Johannes du Plessis
094b2df939
feat: cross-provider model fallback on transient errors (#1281)
When the primary model raises a transient provider error (5xx, 429,
connection/timeout) the request is retried once against a fallback
model from the other provider. Anthropic primaries fall back to
OpenAI and vice versa. Also bumps the SDK max_retries from the
default 2 to 6 so quick blips stay on the primary and keep prompt
caching warm.

Triggered by 529 OverloadedError traces that ended runs silently
with no Slack/Linear/PR reply.
2026-05-08 15:35:13 -07:00
open-swe[bot]
a9331e78e6
fix: harden http_request SSRF guard against DNS rebinding [closes AB-2321] (#1277)
* fix: harden http_request SSRF guard against DNS rebinding [closes AB-2321]

Pin DNS resolution per request hop so urllib3's connection-time lookup
cannot rebind to a private IP after _is_url_safe validated a public one.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: scope DNS pin to urllib3 with reference-counted install

Address review feedback: the previous version permanently overwrote the
process-global socket.getaddrinfo on first use. Now the patch targets
urllib3.util.connection.create_connection (much narrower blast radius),
and is installed/uninstalled via reference count so no global mutation
persists once no http_request calls are in flight.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: forward timeout and socket_options through pinned create_connection

urllib3 calls create_connection with timeout positional and socket_options
as a keyword. The previous wrapper only read kwargs, silently dropping the
caller's connect timeout (so a slow validated IP could hang) and TCP
options like TCP_NODELAY. Accept both positionally and forward them to
the underlying socket.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 22:06:25 +00:00
Johannes du Plessis
b541f6d08a
slack: slim down the trace reply (drop greeting, suppress unfurl) (#1278)
Now that assistant.threads.setStatus carries the "is thinking…"
indicator, the per-run greeting phrase + auto-unfurl on the LangSmith
link are just visual noise stacking on top of it. The trace reply now
posts only `<url|View trace>` + a tip, with link unfurling off so the
smith.langchain.com card no longer appears.

Adds an `unfurl_links`/`unfurl_media` knob to post_slack_thread_reply_with_ts
(default-on to preserve behaviour for every other caller).
2026-05-08 14:51:28 -07:00
open-swe[bot]
148efeb269
feat: support TOKEN_ENCRYPTION_KEY rotation via MultiFernet [closes AB-2323] (#1275)
Threat model T6: Fernet token-encryption key had no rotation path. Rotating
TOKEN_ENCRYPTION_KEY immediately invalidated every github_token_encrypted
value in thread metadata, forcing re-auth or bot-token re-resolution.

Switch to cryptography.fernet.MultiFernet and parse TOKEN_ENCRYPTION_KEY as a
comma- or newline-separated ordered list (most-recent-first). New writes
encrypt under the first key; reads try every key. Deployers can prepend a new
key, let active threads roll over, then drop the old key.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-08 14:29:23 -07:00
Johannes du Plessis
3f9dbb6597
feat: add Slack reaction feedback to LangSmith (#1231)
* feat: add Slack reaction feedback to LangSmith

Record Slack reaction feedback against explicitly mapped LangGraph runs so user ratings are idempotent and tied to the message they reacted to.

* Address review feedback on Slack reaction → LangSmith feedback

- langsmith.py: drop lru_cache on _build_langsmith_feedback_clients so
  rotated keys / late env hydration are picked up; dedupe by (key, url)
  tuple instead of key alone so the same key pointing at different
  endpoints (cloud + self-hosted) builds both clients.
- langsmith.py: treat LangSmithNotFoundError on delete_feedback as
  success — out-of-order or redelivered reaction_removed events would
  otherwise loop forever on Slack's retry policy.
- slack_feedback.py: include channel_id in _feedback_key so the same
  message_ts in two channels can't collide on the same feedback id.
- slack_feedback.py: treat conflicting +/- reactions from one user as
  ambiguous (clear feedback) instead of averaging to a misleading 0.5.
- slack_feedback.py + slack.py + webapp.py: gate reaction handling to
  the user who triggered the run (stored in the slack_run_map mapping
  alongside run_id). Prevents bystanders in shared channels from
  polluting eval feedback.
2026-05-08 14:24:12 -07:00
Brace Sproul
1f8a83d4f8
feat: support allowed repos in addition to allowed orgs for webhook filtering (#1092)
* feat: add ALLOWED_GITHUB_REPOS env var for owner/repo-level webhook filtering

Previously only org-level filtering was supported via ALLOWED_GITHUB_ORGS.
This adds ALLOWED_GITHUB_REPOS for finer-grained control, allowing specific
owner/repo pairs to be allowlisted independently of org membership.

* fix: update test patches for _is_repo_org_allowed rename

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 13:26:30 -07:00
open-swe[bot]
37c194dc6d
feat: give agent self-awareness of its own repo (langchain-ai/open-swe) (#1266)
* feat: tell agent its source code lives at langchain-ai/open-swe

* fix: scope self-reference to only when user talks to the agent about itself

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-08 13:13:35 -07:00
Johannes du Plessis
d38f17ffe8
fix Slack assistant status endpoint (#1272) 2026-05-08 13:07:32 -07:00
Johannes du Plessis
743b2b9ba4
fix: recover from mid-run sandbox death (#1274)
* fix: recover from mid-run sandbox death

Recreate dead sandboxes during tool execution and stop repeated unrecoverable timeout loops with a user-facing notification.

* fix: count repeated sandbox recreations

Treat consecutive sandbox recreations as an unrecovered failure streak so outages cannot loop until the model-call limit.
2026-05-08 12:55:36 -07:00
open-swe[bot]
dc9a0b98da
feat: gate @open-swe mentions on public repos to org members (#1273)
Adds a webhook-level check so only members of $PUBLIC_REPO_ORG_GATE
(e.g. langchain-ai) can trigger Open SWE via mentions or review
requests on public repositories. Private repos remain governed by the
existing org/repo allowlists. Internal bots bypass the gate.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-08 11:38:29 -07:00
Johannes du Plessis
5d1020e8d7
fix: restore Co-authored-by trailer for triggering user (#1271)
The gh-cli migration removed agent/middleware/open_pr.py and
agent/tools/commit_and_open_pr.py — the only callers of
add_user_coauthor_trailer / add_pr_collaboration_note. Since then the
agent has been driving commits and PRs entirely via gh, with no
attribution back to the Slack/Linear/GitHub user who triggered the run.

Resolve the triggering user's identity in get_agent (reusing the
existing authorship helpers) and inject a Collaborative Attribution
section into the system prompt with the exact trailer and PR-body note
to use. The section is only rendered when an identity is resolvable, so
runs without a known triggering user are unchanged.
2026-05-08 10:42:15 -07:00
open-swe[bot]
96f97710ad
feat: add optional Slack Assistants API typing status indicator (#1269)
* feat: add optional Slack Assistants API typing status indicator

Mirrors OpenClaw's pragmatic approach: instead of rebuilding around
assistant_thread_started events, just opt into assistants.threads.setStatus
to show 'is thinking…' while the agent is working, and clear it when
post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so
it can be toggled without touching code.

* fix(slack): drop redundant clear, add status heartbeat across model calls

- Slack auto-clears the typing indicator on bot post; remove the explicit
  assistants.threads.setStatus("") call from post_slack_thread_reply.
- The indicator expires after ~2 minutes; add a before_model middleware
  that refreshes it on every model tick so it stays visible across long
  agent runs. Reuses the existing slack_thread.{channel_id,thread_ts}
  configurable already plumbed for notify_step_limit.
- chat:write is sufficient on the bot token (assistant:write is on the
  way out per Slack docs); no scope or app-config change required.

* feat(slack): contextual status text + rotating loading_messages

- set_slack_assistant_status now accepts an optional loading_messages list
  (capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus
  payload so Slack rotates through them client-side.
- The heartbeat middleware derives a contextual status from the last
  assistant message's tool calls (e.g. "searching the codebase…" after
  grep, "running commands…" after execute), falling back to the default
  "is thinking…" when no tool calls or unknown tool name.
- Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the
  contextual status on each refresh.

* fix slack assistant status lifecycle

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 10:21:55 -07:00
open-swe[bot]
5a845ba99f
feat: add Tip section to Slack trace-reply initial message (#1268)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-08 09:25:06 -07:00
open-swe[bot]
c4d9ae4a67
feat: add idle TTL and delete-after-stop sandbox lifecycle controls (#1265)
* feat: add idle TTL and delete-after-stop sandbox lifecycle controls

* bump langsmith>=0.8.3, lower default idle TTL to 10 min

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-08 00:30:14 -04:00
Johannes du Plessis
bdea1fb6da
feat(open-swe): always anchor findings to a single line (#1264)
* feat(open-swe): always anchor findings to a single line

GitHub renders multi-line review comment ranges as walls of context above
the comment, which buries the point. Match Devin Review's behavior and
always collapse `end_line` to `start_line` so each finding is anchored to
the single most relevant line.

Removes the (now unused) `MAX_FINDING_RANGE_LINES` cap and
`clip_finding_range` helper.

* fix(open-swe): collapse end_line before diff-range check
2026-05-07 20:26:59 -07:00
Johannes du Plessis
089db06135
feat(open-swe): collapse oversized finding ranges to start line (#1263)
Anchoring a finding to a 25-line range (e.g. an entire function) makes
GitHub render the whole block as context above the comment, burying the
review text. Cap ranges at 10 lines and collapse to the start line when
exceeded; small ranges (≤10 lines) still render as a multi-line anchor.
Updated the reviewer prompt to anchor tightly rather than range-select
the whole function.
2026-05-07 19:59:25 -07:00
Johannes du Plessis
0212a60f10
feat(open-swe): cap reviewer suggestions at 4 lines (#1262)
* feat(open-swe): cap reviewer suggestions at 4 lines

Long suggestion blocks read as the reviewer rewriting the code rather
than flagging an issue, which clutters PR comments. Steer the reviewer
toward description-only findings for non-trivial fixes, and enforce the
cap in `add_finding` / `update_finding` so suggestions over 4 lines are
dropped (the finding itself still publishes).

* fix(open-swe): don't clobber prior suggestion on over-cap update

`update_finding` was setting `suggestion=None` whenever clip_suggestion
dropped the input, which silently wiped any existing suggestion on the
finding. Distinguish the three cases: empty string clears, valid value
sets, over-cap value is rejected without touching the stored field.
2026-05-07 19:27:16 -07:00
Johannes du Plessis
7e1746f420
feat(open-swe): trigger reviewer from @open-swe review PR comment (#1259)
* feat(open-swe): trigger reviewer agent from `@open-swe review` PR comment

Mirrors the Slack `@open-swe review` flow on GitHub: a comment containing
`@open-swe review` (optionally followed by a PR URL) on a PR triggers the
reviewer agent. Without a URL it reviews the commenting PR; with a URL it
targets that PR. Works for `issue_comment`, `pull_request_review_comment`,
and `pull_request_review` events, gated by the existing reviewer repo
allowlist and reusing `trigger_pr_review_from_ref`.

* fix(open-swe): require URL after `@open-swe review`, don't swallow trailing text

The previous regex matched any non-whitespace token after `review`, including
across newlines. Comments like `@open-swe review\nthanks!` parsed as
`(True, "thanks!")`, which then failed PR-URL parsing and was silently
dropped — the user got no review and the comment never reached the regular
PR-comment handler.

Restrict the optional URL token to `https?://\S+` so non-URL trailing text
falls through to `process_github_pr_comment` instead of being eaten by the
review-command branch. Adds regression tests for the multiline and
trailing-word cases.
2026-05-07 17:04:35 -07:00
Johannes du Plessis
892347041f
feat(open-swe): post review summary to Slack on first review (#1258)
When a Slack user kicks off a PR review with `@open-swe review <pr-url>`,
the reviewer agent now posts a one-line summary back to the Slack thread
when it finishes — either "No issues found" or "found N potential
issue(s)" with a link to the GitHub review.

The reviewer agent has no Slack tools by design, so the summary is sent
host-side from `publish_review` after the GitHub review POST succeeds.
The Slack channel/thread_ts is persisted on reviewer thread metadata at
trigger time and read back on publish. Re-reviews triggered by push
events stay silent in Slack to avoid noise on the original thread.
2026-05-07 16:11:36 -07:00
Johannes du Plessis
88a55057ea
fix(open-swe): host-format the review summary, drop agent prose (#1257)
The reviewer agent was writing a 1–2 sentence "top-level take" as the
review body, which produced noisy paragraph-style summaries on PRs
("Reviewed the PR. The new ALLOWED_GITHUB_REPOS allowlist…"). Devin's
review comment is just a one-liner (`✅ No Issues Found` or
`**Devin Review** found N potential issue.`), and was preferred in the
internal A/B vs Graphite.

Drop the `summary` parameter from `publish_review`; render a fixed,
host-formatted body in `render_review_body` instead. Update the reviewer
prompt to forbid prose summaries.
2026-05-07 15:41:40 -07:00
Johannes du Plessis
03d9f23645
fix(open-swe): always post a review summary, even with no findings (#1256)
* fix(reviewer): log every push/close early-return so 'silent ignore' is debuggable

Pushes to PRs that haven't had a first review fall through the watch
handler because the reviewer thread doesn't have kind=reviewer set.
Without log lines on the early-return paths, this scenario was
indistinguishable from 'webhook reached the handler at all' in the
hosted log stream.

Now every early-return logs at info or debug:
- info when a real PR exists but the reviewer thread isn't set up
  (with a hint pointing at the trigger paths the user can use)
- info when the repo isn't in the reviewer allowlist
- debug for benign skips (non-branch refs, branch deletions,
  already-reviewed head_sha)

* fix(reviewer): always post a summary review, even with no findings

The publish_review tool gated POSTing on `inline_comments or summary`,
so when the agent called publish_review() with no args on a clean PR
the result returned `success: true` but no GitHub review was posted —
the user got silence instead of a "no issues found" comment.

- Drop the gate so publish_review always POSTs.
- Friendlier no-findings render: `**No issues found.**` when the
  findings list is empty, vs. `**No issues at or above \`<sev>\`
  severity.**` with hidden count when only sub-threshold findings
  exist. Agent summary renders below.
- Prompt now requires the agent to always pass a `summary` so the
  body is meaningful; calls out specifically not to skip on a clean PR.
2026-05-07 22:16:20 +00:00
Johannes du Plessis
378b95266e
feat: implement reviewer findings, publish_review, and watch mode (#1253)
* feat: implement reviewer findings, publish_review, and watch mode

Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:

- Findings as first-class state on the reviewer thread metadata
  (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
  ranges, suggestion text for ```suggestion blocks, github_review_comment_id
  for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
  metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
  future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
  compute_diff_line_set for in-diff validation, extract_diff_hunk for
  caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
  diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
  ranges fail at creation, not at GitHub-publish), `update_finding`,
  `list_findings`, `publish_review`. The reviewer agent's tool list is
  swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
  one POST /reviews call with body + inline comments + ```suggestion blocks,
  per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
  fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
  before the agent's first model call (warm- and cold-path symmetric);
  computed diff and in-diff line set passed via runnable config; system
  prompt rewritten for the single-evolving-findings model, severity ladder,
  in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
  added to supported events. New `process_github_push_event` resolves the
  open PR for the pushed branch, gates on the reviewer thread's `watch`
  flag, builds a re-review configurable, and triggers a run on the same
  canonical thread. `process_github_pr_close` toggles watch on
  closed/reopened. `set_reviewer_thread_metadata` is called on first
  review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
  legacy {file, line, body, severity} shape the judge expects) and passes
  the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
  publish rendering + GraphQL resolve, and watch-mode webhook handlers
  (push triggers re-review only when watching, idempotent on unchanged
  head SHA, PR close disables watch). Updated existing reviewer-webhook
  tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
  into REVIEWER_DESIGN.md.

* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL

Address PR #1253 review findings:

- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
  (`option no-prefix takes no value` — every prep run was failing
  silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
  now uses three-dot `base...head` (the merge-base diff GitHub renders
  on Files-changed) so we don't pick up changes that landed on the base
  branch after the PR diverged. Re-review delta keeps two-dot
  `last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
  `github_review_comment_id`. Without this, watched re-reviews
  re-posted every previously surfaced finding, and only the most-recent
  duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
  `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
  is the canonical endpoint; the old form 404s, so comment ids were
  never stored and watch-mode resolution couldn't run.

Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.

* fix(reviewer): default publish cap from 15 to 4

A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
open-swe[bot]
5a01aff1c7
feat: only post Slack 'Working on it!' on first thread mention (#1250)
* feat: only post Slack 'Working on it!' on first thread mention

* feat: randomize Slack trace reply phrase

Pick from a small list of friendly phrases instead of always saying
'Working on it!' so the bot feels less robotic. Explicit messages (e.g.
'Taking a look...' from PR review path) are unaffected.

* adjust phrases

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-07 12:37:46 -07:00
Johannes du Plessis
bb836b58bc
fix: drop multitask_strategy=enqueue and forward image_urls on slack busy path (#1251)
The Slack mention path already routes mid-run messages through the
store-based queue + check_message_queue_before_model middleware (the
same path Linear uses), so multitask_strategy="enqueue" on runs.create
was a leftover no-op — that line is only reached when is_thread_active
returned False. The busy-path payload was also hardcoding image_urls=[]
which silently dropped any images attached to mid-run Slack mentions;
forward the resolved image_urls so the middleware can rebuild image
blocks like Linear does.
2026-05-07 11:55:14 -07:00
Johannes du Plessis
88d3659d00
start sandbox before proxy refresh (#1249) 2026-05-07 11:26:25 -07:00
Johannes du Plessis
da74342da4
add logging (#1248) 2026-05-07 10:19:53 -07:00
Johannes du Plessis
6ded489751
fix: reuse reviewer thread github token (#1247) 2026-05-06 18:09:39 -07:00
Johannes du Plessis
d060b8a863
fix: reduce GitHub webhook ignore log noise (#1246) 2026-05-07 01:01:33 +00:00
Johannes du Plessis
1319347dd9
feat: route Slack PR review requests (#1245)
* feat: route Slack PR review requests

Add a lightweight Slack review command path that starts the reviewer graph directly and gives the core agent a handoff tool when review requests are misrouted.

* fix: harden Slack PR review routing

* fix: validate Slack PR review URLs

* fix: preserve malformed GitHub review routing
2026-05-06 17:14:43 -07:00
Johannes du Plessis
65f6b4636b
feat(open-swe): remove github_comment and enable reviewer trigger (#1244)
* Use gh for reviewer inline comments

* Trigger reviewer on PR review requests

* Format reviewer webhook test

* Add GitHub repo allowlist

* Separate reviewer repo allowlist

* only implement allowlist for reviewer
2026-05-06 16:14:38 -07:00
Johannes du Plessis
96774f20ae
feat: move github workflows to gh cli (#1238)
* feat: move github workflows to gh cli

Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.

* docker ignore + snapshot and docker image updates

* updated image and instructions

* removing open_pr if needed after agent call
2026-05-04 18:03:53 -07:00
Johannes du Plessis
13f5d8a1c9
fix: preserve existing PR descriptions (#1237)
* fix: preserve existing PR descriptions

* test: update existing PR label expectations
2026-05-04 11:43:32 -07:00
Johannes du Plessis
30530bf4d5
fix: preserve existing PR titles (#1236)
Keep automatic existing-PR updates from retitling pull requests, while documenting the explicit edit path for intentional title changes.
2026-05-04 09:35:08 -07:00
Brace Sproul
3405d145ac
feat: add edit_pull_request tool for editing PR titles/descriptions (#1063)
* feat: add edit_pull_request tool for editing PR titles and descriptions

* fix: patch auth flow in open PR middleware tests

* fix: support app token for editing PRs

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 18:12:41 -07:00
fuyua9
226e6c8263
fix(daytona): make sandbox snapshot configurable (#1220) 2026-05-01 22:51:34 +00:00
langsmith-forge[bot]
bd97678a5e
fix: prevent futile retry loop when commit_and_open_pr fails (#1210)
* fix: prevent futile retry loop when commit_and_open_pr fails with git/API errors

- Root cause: when git checkout or GitHub PR API fails, the tool returned a generic {"success": false} error with no signal that retrying is futile, causing the agent to loop 9-13+ times until hitting the 1000-step recursion limit
- Change: (1) git_checkout_branch now returns (bool, str) so the actual git error output is surfaced in the tool response; (2) checkout and PR creation failures now include "fatal": true and an explicit "Do not retry" message; (3) prompt.py COMMIT_PR_SECTION adds an explicit instruction to stop on fatal errors
- Verified: 109 unit tests pass, no regressions

* fix: skip PR safety net on fatal commit failures

* style(open_pr): ruff-format fatal retry skip condition

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 22:05:22 +00:00
langsmith-forge[bot]
48965a82a6
fix: coerce malformed integer strings in read_file offset/limit params (#1216)
- Root cause: LLM occasionally generates strings like '1, 80' or '170, "limit": 60'
  for integer fields, causing a Pydantic ValidationError and wasting an LLM turn
- Change: add SanitizeToolInputsMiddleware in agent/middleware/sanitize_tool_inputs.py
  that extracts the leading integer from any string value in offset/limit before
  the call reaches Pydantic validation; registered before ToolErrorMiddleware in server.py
- Verified: 14 unit tests covering all three production trace patterns pass

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:29:48 -07:00
langsmith-forge[bot]
e5bc27a0ad
fix: notify users via Slack when agent hits model call step limit (#1204)
* fix: notify users via Slack when agent hits model call step limit

- Root cause: GraphRecursionError at 1000 steps bypassed all @after_agent
  middleware including open_pr_if_needed, leaving users with no notification
- Change: Added ModelCallLimitMiddleware(run_limit=60) to intercept gracefully
  before the hard recursion limit, and added notify_step_limit_reached
  @after_agent middleware to post a Slack thread reply when the limit fires
- Verified: 107 existing tests pass, no regressions

* fix: harden step-limit Slack notification

Ensure the step-limit notification runs after the PR safety net and cover the new middleware behavior with focused unit tests.

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:24:25 -07:00
langsmith-forge[bot]
8095f93364
fix: update existing PR title and body when create returns 422 (#1163)
- Root cause: create_github_pr found an existing PR on 422 but never
  PATCHed it, so callers like commit_and_open_pr could not update the
  PR body (e.g. adding "Closes AB-1159") on subsequent invocations.
- Change: after _find_existing_pr succeeds, call new _update_github_pr
  helper which PATCHes /repos/{owner}/{repo}/pulls/{number} with the
  requested title and body before returning pr_existing=True.
- Verified: self-evident API call addition; proof in production traces.

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-01 14:02:39 -07:00
langsmith-forge[bot]
15a7268d30
fix: return actionable error when push fails with workflows scope error (#1153)
* fix: return actionable error when git push fails due to workflows scope

- Root cause: push failure with "workflows scope" error returned a generic
  "Git push failed: ..." message, causing the agent to retry 10+ times
- Change: detect "workflows" + "scope" in push output and return a clear
  message instructing the agent to remove .github/workflows/ file changes
- Verified: unit tests cover both the workflow-scope path and the
  non-workflow path to prevent regressions

* fix: detect github workflow permission push failures

* style: ruff-format checkout line in commit_and_open_pr

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 13:09:54 -07:00
langsmith-forge[bot]
b25222f1f9
fix: recover existing PR after httpx.HTTPError in create_github_pr (#1148)
- Root cause: httpx.HTTPError handler returned (None, None, False) without
  checking if the PR was already created on GitHub before the network error
- Change: added _find_existing_pr fallback in except block in agent/utils/github.py
- Verified: 5 production traces showed false failures where PR existed (pr_existing=True on retry)

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 13:00:23 -07:00
langsmith-forge-dev[bot]
c61d8b0376
fix: stop agent retrying commit_and_open_pr on 403 permission denied (#1123)
* fix: stop agent retrying commit_and_open_pr on 403 permission denied

Detect 403/permission-denied push failures in commit_and_open_pr and
return a PERMANENT_FAILURE message so the LLM stops retrying. Also add
prompt-level guidance to the COMMIT_PR_SECTION reinforcing this. Add
unit tests covering both the 403 and non-403 push failure paths.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: stop safety net retrying permanent push failures

Skip the after-agent PR fallback when commit_and_open_pr reports a permanent GitHub push authorization failure, while preserving fallback behavior for recoverable failures.

---------

Co-authored-by: Claude Agent <agent@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 19:37:53 +00:00
langsmith-forge-dev[bot]
4060933ce4
feat: add github CI check run tools for shepherding CI (#1121)
* feat: add github CI check run tools for shepherding CI

Add get_pr_check_runs and rerun_failed_check_runs tools that authenticate
using the GitHub App installation token so the agent can query and retry
CI status on private repos without relying on GH_TOKEN or unauthenticated
http_request calls.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: handle paginated GitHub CI results

* refactor(github_ci): address review feedback

- Rename rerun_failed_check_runs -> rerun_failed_workflow_runs and clarify
  in docstrings that the tool only retries GitHub Actions workflow runs
  (not third-party CI checks surfaced by get_pr_check_runs).
- Skip action_required workflow runs when rerunning; those need manual
  approval, not a rerun.
- Run rerun-failed-jobs requests concurrently via asyncio.gather instead
  of sequentially.
- Fix latent pagination bug in _fetch_paginated_items where caller-supplied
  params could overwrite per_page/page and break the end-of-pagination
  check; reserved keys now always win and the threshold uses a PER_PAGE
  constant.
- Set an explicit 30s httpx timeout so a hung GitHub call cannot stall
  the agent loop.
- Restore alphabetical ordering of tools in agent/tools/__init__.py.
- Add tests for: a 500 surfaced on a later pagination page, and
  action_required runs being filtered out of rerun candidates.

---------

Co-authored-by: Claude Agent <agent@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 12:19:04 -07:00
langsmith-forge-dev[bot]
b5ed2a6b8b
fix: safety net middleware never fires due to key-existence check (#1051)
* fix: safety net middleware always skipped due to key-existence check

The open_pr_if_needed after-agent middleware checked `if 'success' in pr_payload`
which evaluates True for BOTH success and failure responses from commit_and_open_pr
(all responses include the 'success' key). This meant the safety net never fired.

Fix: use `pr_payload.get('success')` to check the VALUE instead of key existence.

Evidence: 6+ production traces in last 24h where commit_and_open_pr returned
success=False but the safety net silently skipped (non-fast-forward push failures,
missing GitHub token, workflow permission errors, API 500 errors).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* update

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Palash Shah <palash@langchain.dev>
2026-04-30 17:22:11 -07:00
langsmith-forge-dev[bot]
d63780a77f
fix: add get_pr_review_comments tool to fetch PR review comments with auth (#1043)
* feat: add get_pr_review_comments tool for authenticated GitHub API access

The agent was asking users to paste PR review comments because it had no
tool to fetch them with auth. This adds get_pr_review_comments, which uses
the GitHub App installation token to fetch all three comment types (thread
comments, inline review comments, review submissions) from private repos.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* u

* u

---------

Co-authored-by: Forge Agent <agent@forge.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Palash Shah <palash@langchain.dev>
Co-authored-by: Palash Shah <35114859+Palashio@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-04-30 17:01:13 -07:00
Aran Yogesh
920a8a7624
feat: open PRs under user's name and add OpenSWE label (#1215)
* feat: open PRs under user's name and add OpenSWE label

* feat: use user token for PR authorship, add OpenSWE label, and consolidate fallback logic

* linting

* fix: address review nits for PR authorship and labeling

Fix docstring casing, add debug logging for 422 existing-PR search
fallback, tighten test type annotations, and add missing HTTPError
fallback test.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-04-29 17:34:00 -07:00