Commit graph

44 commits

Author SHA1 Message Date
open-swe[bot]
a9331e78e6
fix: harden http_request SSRF guard against DNS rebinding [closes AB-2321] (#1277)
* fix: harden http_request SSRF guard against DNS rebinding [closes AB-2321]

Pin DNS resolution per request hop so urllib3's connection-time lookup
cannot rebind to a private IP after _is_url_safe validated a public one.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: scope DNS pin to urllib3 with reference-counted install

Address review feedback: the previous version permanently overwrote the
process-global socket.getaddrinfo on first use. Now the patch targets
urllib3.util.connection.create_connection (much narrower blast radius),
and is installed/uninstalled via reference count so no global mutation
persists once no http_request calls are in flight.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: forward timeout and socket_options through pinned create_connection

urllib3 calls create_connection with timeout positional and socket_options
as a keyword. The previous wrapper only read kwargs, silently dropping the
caller's connect timeout (so a slow validated IP could hang) and TCP
options like TCP_NODELAY. Accept both positionally and forward them to
the underlying socket.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 22:06:25 +00:00
Johannes du Plessis
3f9dbb6597
feat: add Slack reaction feedback to LangSmith (#1231)
* feat: add Slack reaction feedback to LangSmith

Record Slack reaction feedback against explicitly mapped LangGraph runs so user ratings are idempotent and tied to the message they reacted to.

* Address review feedback on Slack reaction → LangSmith feedback

- langsmith.py: drop lru_cache on _build_langsmith_feedback_clients so
  rotated keys / late env hydration are picked up; dedupe by (key, url)
  tuple instead of key alone so the same key pointing at different
  endpoints (cloud + self-hosted) builds both clients.
- langsmith.py: treat LangSmithNotFoundError on delete_feedback as
  success — out-of-order or redelivered reaction_removed events would
  otherwise loop forever on Slack's retry policy.
- slack_feedback.py: include channel_id in _feedback_key so the same
  message_ts in two channels can't collide on the same feedback id.
- slack_feedback.py: treat conflicting +/- reactions from one user as
  ambiguous (clear feedback) instead of averaging to a misleading 0.5.
- slack_feedback.py + slack.py + webapp.py: gate reaction handling to
  the user who triggered the run (stored in the slack_run_map mapping
  alongside run_id). Prevents bystanders in shared channels from
  polluting eval feedback.
2026-05-08 14:24:12 -07:00
Johannes du Plessis
bdea1fb6da
feat(open-swe): always anchor findings to a single line (#1264)
* feat(open-swe): always anchor findings to a single line

GitHub renders multi-line review comment ranges as walls of context above
the comment, which buries the point. Match Devin Review's behavior and
always collapse `end_line` to `start_line` so each finding is anchored to
the single most relevant line.

Removes the (now unused) `MAX_FINDING_RANGE_LINES` cap and
`clip_finding_range` helper.

* fix(open-swe): collapse end_line before diff-range check
2026-05-07 20:26:59 -07:00
Johannes du Plessis
089db06135
feat(open-swe): collapse oversized finding ranges to start line (#1263)
Anchoring a finding to a 25-line range (e.g. an entire function) makes
GitHub render the whole block as context above the comment, burying the
review text. Cap ranges at 10 lines and collapse to the start line when
exceeded; small ranges (≤10 lines) still render as a multi-line anchor.
Updated the reviewer prompt to anchor tightly rather than range-select
the whole function.
2026-05-07 19:59:25 -07:00
Johannes du Plessis
0212a60f10
feat(open-swe): cap reviewer suggestions at 4 lines (#1262)
* feat(open-swe): cap reviewer suggestions at 4 lines

Long suggestion blocks read as the reviewer rewriting the code rather
than flagging an issue, which clutters PR comments. Steer the reviewer
toward description-only findings for non-trivial fixes, and enforce the
cap in `add_finding` / `update_finding` so suggestions over 4 lines are
dropped (the finding itself still publishes).

* fix(open-swe): don't clobber prior suggestion on over-cap update

`update_finding` was setting `suggestion=None` whenever clip_suggestion
dropped the input, which silently wiped any existing suggestion on the
finding. Distinguish the three cases: empty string clears, valid value
sets, over-cap value is rejected without touching the stored field.
2026-05-07 19:27:16 -07:00
Johannes du Plessis
892347041f
feat(open-swe): post review summary to Slack on first review (#1258)
When a Slack user kicks off a PR review with `@open-swe review <pr-url>`,
the reviewer agent now posts a one-line summary back to the Slack thread
when it finishes — either "No issues found" or "found N potential
issue(s)" with a link to the GitHub review.

The reviewer agent has no Slack tools by design, so the summary is sent
host-side from `publish_review` after the GitHub review POST succeeds.
The Slack channel/thread_ts is persisted on reviewer thread metadata at
trigger time and read back on publish. Re-reviews triggered by push
events stay silent in Slack to avoid noise on the original thread.
2026-05-07 16:11:36 -07:00
Johannes du Plessis
88a55057ea
fix(open-swe): host-format the review summary, drop agent prose (#1257)
The reviewer agent was writing a 1–2 sentence "top-level take" as the
review body, which produced noisy paragraph-style summaries on PRs
("Reviewed the PR. The new ALLOWED_GITHUB_REPOS allowlist…"). Devin's
review comment is just a one-liner (`✅ No Issues Found` or
`**Devin Review** found N potential issue.`), and was preferred in the
internal A/B vs Graphite.

Drop the `summary` parameter from `publish_review`; render a fixed,
host-formatted body in `render_review_body` instead. Update the reviewer
prompt to forbid prose summaries.
2026-05-07 15:41:40 -07:00
Johannes du Plessis
03d9f23645
fix(open-swe): always post a review summary, even with no findings (#1256)
* fix(reviewer): log every push/close early-return so 'silent ignore' is debuggable

Pushes to PRs that haven't had a first review fall through the watch
handler because the reviewer thread doesn't have kind=reviewer set.
Without log lines on the early-return paths, this scenario was
indistinguishable from 'webhook reached the handler at all' in the
hosted log stream.

Now every early-return logs at info or debug:
- info when a real PR exists but the reviewer thread isn't set up
  (with a hint pointing at the trigger paths the user can use)
- info when the repo isn't in the reviewer allowlist
- debug for benign skips (non-branch refs, branch deletions,
  already-reviewed head_sha)

* fix(reviewer): always post a summary review, even with no findings

The publish_review tool gated POSTing on `inline_comments or summary`,
so when the agent called publish_review() with no args on a clean PR
the result returned `success: true` but no GitHub review was posted —
the user got silence instead of a "no issues found" comment.

- Drop the gate so publish_review always POSTs.
- Friendlier no-findings render: `**No issues found.**` when the
  findings list is empty, vs. `**No issues at or above \`<sev>\`
  severity.**` with hidden count when only sub-threshold findings
  exist. Agent summary renders below.
- Prompt now requires the agent to always pass a `summary` so the
  body is meaningful; calls out specifically not to skip on a clean PR.
2026-05-07 22:16:20 +00:00
Johannes du Plessis
378b95266e
feat: implement reviewer findings, publish_review, and watch mode (#1253)
* feat: implement reviewer findings, publish_review, and watch mode

Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:

- Findings as first-class state on the reviewer thread metadata
  (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
  ranges, suggestion text for ```suggestion blocks, github_review_comment_id
  for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
  metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
  future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
  compute_diff_line_set for in-diff validation, extract_diff_hunk for
  caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
  diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
  ranges fail at creation, not at GitHub-publish), `update_finding`,
  `list_findings`, `publish_review`. The reviewer agent's tool list is
  swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
  one POST /reviews call with body + inline comments + ```suggestion blocks,
  per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
  fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
  before the agent's first model call (warm- and cold-path symmetric);
  computed diff and in-diff line set passed via runnable config; system
  prompt rewritten for the single-evolving-findings model, severity ladder,
  in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
  added to supported events. New `process_github_push_event` resolves the
  open PR for the pushed branch, gates on the reviewer thread's `watch`
  flag, builds a re-review configurable, and triggers a run on the same
  canonical thread. `process_github_pr_close` toggles watch on
  closed/reopened. `set_reviewer_thread_metadata` is called on first
  review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
  legacy {file, line, body, severity} shape the judge expects) and passes
  the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
  publish rendering + GraphQL resolve, and watch-mode webhook handlers
  (push triggers re-review only when watching, idempotent on unchanged
  head SHA, PR close disables watch). Updated existing reviewer-webhook
  tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
  into REVIEWER_DESIGN.md.

* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL

Address PR #1253 review findings:

- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
  (`option no-prefix takes no value` — every prep run was failing
  silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
  now uses three-dot `base...head` (the merge-base diff GitHub renders
  on Files-changed) so we don't pick up changes that landed on the base
  branch after the PR diverged. Re-review delta keeps two-dot
  `last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
  `github_review_comment_id`. Without this, watched re-reviews
  re-posted every previously surfaced finding, and only the most-recent
  duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
  `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
  is the canonical endpoint; the old form 404s, so comment ids were
  never stored and watch-mode resolution couldn't run.

Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.

* fix(reviewer): default publish cap from 15 to 4

A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
open-swe[bot]
5a01aff1c7
feat: only post Slack 'Working on it!' on first thread mention (#1250)
* feat: only post Slack 'Working on it!' on first thread mention

* feat: randomize Slack trace reply phrase

Pick from a small list of friendly phrases instead of always saying
'Working on it!' so the bot feels less robotic. Explicit messages (e.g.
'Taking a look...' from PR review path) are unaffected.

* adjust phrases

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-07 12:37:46 -07:00
Johannes du Plessis
1319347dd9
feat: route Slack PR review requests (#1245)
* feat: route Slack PR review requests

Add a lightweight Slack review command path that starts the reviewer graph directly and gives the core agent a handoff tool when review requests are misrouted.

* fix: harden Slack PR review routing

* fix: validate Slack PR review URLs

* fix: preserve malformed GitHub review routing
2026-05-06 17:14:43 -07:00
Johannes du Plessis
65f6b4636b
feat(open-swe): remove github_comment and enable reviewer trigger (#1244)
* Use gh for reviewer inline comments

* Trigger reviewer on PR review requests

* Format reviewer webhook test

* Add GitHub repo allowlist

* Separate reviewer repo allowlist

* only implement allowlist for reviewer
2026-05-06 16:14:38 -07:00
Johannes du Plessis
ace71b0fd0
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring

- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
  alongside the main `agent` graph. Reuses the same sandbox lifecycle,
  GH proxy auth, and middleware primitives from `agent.server`, but with
  a narrower tool set, a reviewer-specific system prompt, no
  commit/push, and the `task` (subagent) tool stripped via
  `_ToolExclusionMiddleware` so review stays in one context.

- New `github_comment` tool: agents call it once per issue with
  `(file, line, body, severity)` and the eval scores those calls
  against golden comments.

- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
  *not* on the reviewer's stack — that middleware exists to enforce the
  main agent's "always finalize via Slack/Linear/PR" contract, which
  the reviewer doesn't have. The main agent's behavior is unchanged.

- `evals/reviewer/target.py`: send PR info as a user message, extract
  every `github_comment` tool call (multiple expected per review) into
  the run output.

- `evals/reviewer/judge.py`: per-example evaluator now returns a list
  of metrics under `{"results": [...]}` so LangSmith averages each
  numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
  the UI. Dropped the broken `aggregate_pr` summary evaluator that
  reached for an attribute that doesn't exist on `RunTree`.

- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
  `client.list_examples(limit=N)` since `aevaluate` doesn't accept
  `max_examples`.

- Makefile: `dev` and `run` targets now use `uv run` so they work
  without an activated venv.

* resolve comments

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
Johannes du Plessis
96774f20ae
feat: move github workflows to gh cli (#1238)
* feat: move github workflows to gh cli

Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.

* docker ignore + snapshot and docker image updates

* updated image and instructions

* removing open_pr if needed after agent call
2026-05-04 18:03:53 -07:00
Johannes du Plessis
13f5d8a1c9
fix: preserve existing PR descriptions (#1237)
* fix: preserve existing PR descriptions

* test: update existing PR label expectations
2026-05-04 11:43:32 -07:00
Brace Sproul
3405d145ac
feat: add edit_pull_request tool for editing PR titles/descriptions (#1063)
* feat: add edit_pull_request tool for editing PR titles and descriptions

* fix: patch auth flow in open PR middleware tests

* fix: support app token for editing PRs

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 18:12:41 -07:00
open-swe[bot]
b829cef97a
feat: add release note section to PR body template (#1202)
* feat: add release note section to PR body template

The langchainplus repo PR template includes a ## Release Note section,
but the agent hardcoded PR body template only had ## Description and
## Test Plan. This caused the release note section to be dropped from
agent-created PRs.

* docs: align PR body argument docs

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 15:32:30 -07:00
langsmith-forge[bot]
bd97678a5e
fix: prevent futile retry loop when commit_and_open_pr fails (#1210)
* fix: prevent futile retry loop when commit_and_open_pr fails with git/API errors

- Root cause: when git checkout or GitHub PR API fails, the tool returned a generic {"success": false} error with no signal that retrying is futile, causing the agent to loop 9-13+ times until hitting the 1000-step recursion limit
- Change: (1) git_checkout_branch now returns (bool, str) so the actual git error output is surfaced in the tool response; (2) checkout and PR creation failures now include "fatal": true and an explicit "Do not retry" message; (3) prompt.py COMMIT_PR_SECTION adds an explicit instruction to stop on fatal errors
- Verified: 109 unit tests pass, no regressions

* fix: skip PR safety net on fatal commit failures

* style(open_pr): ruff-format fatal retry skip condition

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 22:05:22 +00:00
langsmith-forge[bot]
0464126c83
fix: warn against using http_request for GitHub PR creation (#1203)
* fix: warn against using http_request for GitHub PR creation

- Root cause: http_request tool description and system prompt lacked explicit guidance not to use it for GitHub PR operations, causing the agent to fall back to it and receive 401 Unauthorized responses
- Change: added clear warnings to both the http_request docstring and the TOOL_USAGE_SECTION in prompt.py directing agents to use commit_and_open_pr instead
- Verified: docstring and prompt changes are minimal and scoped

* fix: clarify http_request PR guidance

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:14:44 -07:00
langsmith-forge[bot]
15a7268d30
fix: return actionable error when push fails with workflows scope error (#1153)
* fix: return actionable error when git push fails due to workflows scope

- Root cause: push failure with "workflows scope" error returned a generic
  "Git push failed: ..." message, causing the agent to retry 10+ times
- Change: detect "workflows" + "scope" in push output and return a clear
  message instructing the agent to remove .github/workflows/ file changes
- Verified: unit tests cover both the workflow-scope path and the
  non-workflow path to prevent regressions

* fix: detect github workflow permission push failures

* style: ruff-format checkout line in commit_and_open_pr

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 13:09:54 -07:00
langsmith-forge-dev[bot]
c61d8b0376
fix: stop agent retrying commit_and_open_pr on 403 permission denied (#1123)
* fix: stop agent retrying commit_and_open_pr on 403 permission denied

Detect 403/permission-denied push failures in commit_and_open_pr and
return a PERMANENT_FAILURE message so the LLM stops retrying. Also add
prompt-level guidance to the COMMIT_PR_SECTION reinforcing this. Add
unit tests covering both the 403 and non-403 push failure paths.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: stop safety net retrying permanent push failures

Skip the after-agent PR fallback when commit_and_open_pr reports a permanent GitHub push authorization failure, while preserving fallback behavior for recoverable failures.

---------

Co-authored-by: Claude Agent <agent@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 19:37:53 +00:00
langsmith-forge-dev[bot]
4060933ce4
feat: add github CI check run tools for shepherding CI (#1121)
* feat: add github CI check run tools for shepherding CI

Add get_pr_check_runs and rerun_failed_check_runs tools that authenticate
using the GitHub App installation token so the agent can query and retry
CI status on private repos without relying on GH_TOKEN or unauthenticated
http_request calls.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: handle paginated GitHub CI results

* refactor(github_ci): address review feedback

- Rename rerun_failed_check_runs -> rerun_failed_workflow_runs and clarify
  in docstrings that the tool only retries GitHub Actions workflow runs
  (not third-party CI checks surfaced by get_pr_check_runs).
- Skip action_required workflow runs when rerunning; those need manual
  approval, not a rerun.
- Run rerun-failed-jobs requests concurrently via asyncio.gather instead
  of sequentially.
- Fix latent pagination bug in _fetch_paginated_items where caller-supplied
  params could overwrite per_page/page and break the end-of-pagination
  check; reserved keys now always win and the threshold uses a PER_PAGE
  constant.
- Set an explicit 30s httpx timeout so a hung GitHub call cannot stall
  the agent loop.
- Restore alphabetical ordering of tools in agent/tools/__init__.py.
- Add tests for: a 500 surfaced on a later pagination page, and
  action_required runs being filtered out of rerun candidates.

---------

Co-authored-by: Claude Agent <agent@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 12:19:04 -07:00
langsmith-forge-dev[bot]
d63780a77f
fix: add get_pr_review_comments tool to fetch PR review comments with auth (#1043)
* feat: add get_pr_review_comments tool for authenticated GitHub API access

The agent was asking users to paste PR review comments because it had no
tool to fetch them with auth. This adds get_pr_review_comments, which uses
the GitHub App installation token to fetch all three comment types (thread
comments, inline review comments, review submissions) from private repos.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* u

* u

---------

Co-authored-by: Forge Agent <agent@forge.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Palash Shah <palash@langchain.dev>
Co-authored-by: Palash Shah <35114859+Palashio@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-04-30 17:01:13 -07:00
Aran Yogesh
b6ea229a46
feat: give agent ability to read cross-posted Slack message links [closes OPE-37] (#1200)
* feat: give agent ability to read cross-posted Slack message links [close OPE-37]

* refactor: clean up Slack link resolution code

* linting

* refactor: address PR review feedback for Slack link resolution

* linting

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-04-29 17:42:27 -07:00
Aran Yogesh
920a8a7624
feat: open PRs under user's name and add OpenSWE label (#1215)
* feat: open PRs under user's name and add OpenSWE label

* feat: use user token for PR authorship, add OpenSWE label, and consolidate fallback logic

* linting

* fix: address review nits for PR authorship and labeling

Fix docstring casing, add debug logging for 422 existing-PR search
fallback, tighten test type annotations, and add missing HTTPError
fallback test.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-04-29 17:34:00 -07:00
Aran Yogesh
a3a40a1bec
Revert "fix: Auto assign PRs to creator (#1211)" (#1214)
This reverts commit e9b94ac8ba.
2026-04-22 13:46:44 -07:00
Brace Sproul
e9b94ac8ba
fix: Auto assign PRs to creator (#1211)
* fix: Auto assign PRs to creator

* cr
2026-04-21 14:13:55 -07:00
Aran Yogesh
2039fe6660
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159)
* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* removing logger.info

* formatting and linting

* fix: resolve lint errors in server.py (imports, unused vars, undefined names)

* feat: use opaque proxy headers for GitHub auth in sandbox

* linting formatting and test changes

* linting

* Delete .claude directory

* Delete tests/evals directory

* fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests

* fix: restore authorship, branch_name support, and installation token for PR creation

* linitng

* fix: move installation token fetch before commit, clean up dead proxy validation code

* feat: stop auto-cloning and let agent manage repo setup [closes OPE-21]

* feat: stop auto-cloning and let agent manage repo setup [closes OPE-21]

* fix: address review feedback — restore agents_md, add git user config, lint fixes

* fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config

* fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config

* linting

* linting

* feat: add installation token auth to list_repos GitHub API call

* agents.md update

* linting

* fix: address PR review feedback — shell precedence bug in prompt, remove dead code

* linting

* Apply suggestion from @bracesproul

Co-authored-by: Brace Sproul <braceasproul@gmail.com>

* Apply suggestion from @bracesproul

Co-authored-by: Brace Sproul <braceasproul@gmail.com>

* fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block

* fix:Extract check_or_recreate_sandbox utility from inline sandbox health check

* fix: address PR review feedback — async list_repos, restore template name, fix prompt colon

* fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox

* linting

* yogesh/ope-21-stop-auto-cloning

* Update agent/tools/list_repos.py

Co-authored-by: Brace Sproul <braceasproul@gmail.com>

* Update agent/prompt.py

Co-authored-by: Brace Sproul <braceasproul@gmail.com>

* feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo

* linting

* feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check

* feat: support listing repos for personal user accounts via is_organization flag

---------

Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
Aran Yogesh
67c782c295
fix: block open-swe from approving PRs (#1177)
* fix: block open-swe from approving PRs

* linting

* fix: add case-insensitive APPROVE guard and unit tests
2026-04-09 11:45:08 -07:00
Aran Yogesh
4d4f5fbfc7
fix: proxy config restored the branch yogesh/GitHub auth proxy (#1173)
* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* removing logger.info

* formatting and linting

* fix: resolve lint errors in server.py (imports, unused vars, undefined names)

* feat: use opaque proxy headers for GitHub auth in sandbox

* linting formatting and test changes

* linting

* Delete .claude directory

* Delete tests/evals directory

* fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests

* fix: restore authorship, branch_name support, and installation token for PR creation

* linitng

* fix: move installation token fetch before commit, clean up dead proxy validation code

* fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config

* fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config

* linting

* linting

* fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox
2026-04-08 15:02:52 -07:00
Aran Yogesh
9aba4d0545
Revert "feat: authenticate git operations via sandbox proxy instead of creden…" (#1170)
This reverts commit 6305e13dc6.
2026-04-07 18:45:33 -07:00
Aran Yogesh
6305e13dc6
feat: authenticate git operations via sandbox proxy instead of credential files [closes: OPE-20] (#1070)
* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* removing logger.info

* formatting and linting

* fix: resolve lint errors in server.py (imports, unused vars, undefined names)

* feat: use opaque proxy headers for GitHub auth in sandbox

* linting formatting and test changes

* linting

* Delete .claude directory

* Delete tests/evals directory

* fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests

* fix: restore authorship, branch_name support, and installation token for PR creation

* linitng

* fix: move installation token fetch before commit, clean up dead proxy validation code

* fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config

* fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config

* linting

* linting

* fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox
2026-04-07 16:41:48 -07:00
Brace Sproul
fd8e6d98ee
fix: Shell injection and ssrf issues (#1155)
* fix: Shell injection and ssrf issues

* cr
2026-04-01 12:37:35 -07:00
Aran Yogesh
86307affe4
fix: use GitHub App installation token for PR creation instead of user token (#1149)
* fix: use GitHub App installation token for PR creation instead of user token

* fix: move installation token fetch after no-changes check to avoid unnecessary API call
2026-03-30 12:47:57 -07:00
Aran Yogesh
7d1004ad66
fix: add Exa web search tool [closes: OPE-29] (#1133)
* fix: add Exa web search tool

* chore: update uv.lock for exa-py dependency

* linting

* chore: remove web_search from system prompt

* chore: drop search_type and category params from web_search
2026-03-25 16:24:31 -07:00
Brace Sproul
87968ab813
fix: Add scripts for getting usage, fix dual committing (#1131) 2026-03-25 13:39:33 -07:00
Brace Sproul
e1de58c584
fix: convert @Name(USER_ID) mentions to Slack's <@USER_ID> format in thread replies (#1126)
The agent sees users formatted as @Name(USER_ID) in conversation context and
reproduces that pattern in replies, but Slack requires <@USER_ID> for real
mentions. This adds automatic conversion and updates prompt instructions.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-03-25 11:51:11 -07:00
Brace Sproul
267b2fd011
feat: add github pr review tools [closes OPE-24] (#1104)
* Add GitHub PR review tools (list, get, create, update, dismiss, submit, list comments) and bind them to the agent

* format n lint

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Aran Yogesh <yogesh.mahendran@langchain.dev>
2026-03-23 21:37:54 +00:00
Brace Sproul
91348aeb7c
feat: add linear tools for listing teams, get/create/update/delete issues, and get issue comments (#1105)
Adds 6 new agent tools backed by Linear's GraphQL API, with a shared
_graphql_request helper to reduce boilerplate. Refactors existing
comment_on_linear_issue to use the same helper.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-03-23 14:32:44 -07:00
Aran Yogesh
150ff6f0c8
fix: allow open-swe to be triggered on any GitHuh branch (#1100)
* fix: allow open-swe to be triggered on any GitHuh branch

* formmatting
2026-03-20 10:53:33 -07:00
Brace Sproul
d7d9bc5179
fix: better custom backend support (#1071)
* fix: better custom backend support

* cr
2026-03-17 11:55:36 -07:00
Brace Sproul
3e73447920
fix: better docs, github auth (#1055)
* fix: Update github auth methods

* cr
2026-03-11 22:57:56 -08:00
Aran Yogesh
b735afb445
feat: add GitHub PR comment trigger and reply support [closes OPE-1] (#1014)
* feat: add GitHub PR comment trigger and reply support

* refactor: improve readability of GitHub integration

* linting

* fix: resolve github token from thread metadata and improve PR trigger flow

* fix: fall back to OAuth for GitHub webhook when no token in thread metadata

* give me commit message github integeration working without a breaking

* auth.py refactor

* fix: validate cached GitHub token before use to handle expiry

* liniting

* ci unitest formatting

* feat: post PR comments as GitHub App bot instead of user OAuth token

* resolved comments

* slack resolveed comments

* Refactor docstring and comments in get_slack_repo_config

Removed unnecessary comments and cleaned up docstring formatting.

* cr

* cr

* cr

* cr

---------

Co-authored-by: bracesproul <braceasproul@gmail.com>
2026-03-09 17:14:13 -07:00
Brace Sproul
bd52e5e09d
chore: Drop monorepo (#1029)
* chore: Drop monorepo

* cr
2026-03-06 16:10:34 -08:00