Commit graph

22 commits

Author SHA1 Message Date
Johannes du Plessis
ace71b0fd0
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring

- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
  alongside the main `agent` graph. Reuses the same sandbox lifecycle,
  GH proxy auth, and middleware primitives from `agent.server`, but with
  a narrower tool set, a reviewer-specific system prompt, no
  commit/push, and the `task` (subagent) tool stripped via
  `_ToolExclusionMiddleware` so review stays in one context.

- New `github_comment` tool: agents call it once per issue with
  `(file, line, body, severity)` and the eval scores those calls
  against golden comments.

- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
  *not* on the reviewer's stack — that middleware exists to enforce the
  main agent's "always finalize via Slack/Linear/PR" contract, which
  the reviewer doesn't have. The main agent's behavior is unchanged.

- `evals/reviewer/target.py`: send PR info as a user message, extract
  every `github_comment` tool call (multiple expected per review) into
  the run output.

- `evals/reviewer/judge.py`: per-example evaluator now returns a list
  of metrics under `{"results": [...]}` so LangSmith averages each
  numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
  the UI. Dropped the broken `aggregate_pr` summary evaluator that
  reached for an attribute that doesn't exist on `RunTree`.

- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
  `client.list_examples(limit=N)` since `aevaluate` doesn't accept
  `max_examples`.

- Makefile: `dev` and `run` targets now use `uv run` so they work
  without an activated venv.

* resolve comments

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
Johannes du Plessis
96774f20ae
feat: move github workflows to gh cli (#1238)
* feat: move github workflows to gh cli

Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.

* docker ignore + snapshot and docker image updates

* updated image and instructions

* removing open_pr if needed after agent call
2026-05-04 18:03:53 -07:00
langsmith-forge[bot]
bd97678a5e
fix: prevent futile retry loop when commit_and_open_pr fails (#1210)
* fix: prevent futile retry loop when commit_and_open_pr fails with git/API errors

- Root cause: when git checkout or GitHub PR API fails, the tool returned a generic {"success": false} error with no signal that retrying is futile, causing the agent to loop 9-13+ times until hitting the 1000-step recursion limit
- Change: (1) git_checkout_branch now returns (bool, str) so the actual git error output is surfaced in the tool response; (2) checkout and PR creation failures now include "fatal": true and an explicit "Do not retry" message; (3) prompt.py COMMIT_PR_SECTION adds an explicit instruction to stop on fatal errors
- Verified: 109 unit tests pass, no regressions

* fix: skip PR safety net on fatal commit failures

* style(open_pr): ruff-format fatal retry skip condition

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 22:05:22 +00:00
langsmith-forge[bot]
48965a82a6
fix: coerce malformed integer strings in read_file offset/limit params (#1216)
- Root cause: LLM occasionally generates strings like '1, 80' or '170, "limit": 60'
  for integer fields, causing a Pydantic ValidationError and wasting an LLM turn
- Change: add SanitizeToolInputsMiddleware in agent/middleware/sanitize_tool_inputs.py
  that extracts the leading integer from any string value in offset/limit before
  the call reaches Pydantic validation; registered before ToolErrorMiddleware in server.py
- Verified: 14 unit tests covering all three production trace patterns pass

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:29:48 -07:00
langsmith-forge[bot]
e5bc27a0ad
fix: notify users via Slack when agent hits model call step limit (#1204)
* fix: notify users via Slack when agent hits model call step limit

- Root cause: GraphRecursionError at 1000 steps bypassed all @after_agent
  middleware including open_pr_if_needed, leaving users with no notification
- Change: Added ModelCallLimitMiddleware(run_limit=60) to intercept gracefully
  before the hard recursion limit, and added notify_step_limit_reached
  @after_agent middleware to post a Slack thread reply when the limit fires
- Verified: 107 existing tests pass, no regressions

* fix: harden step-limit Slack notification

Ensure the step-limit notification runs after the PR safety net and cover the new middleware behavior with focused unit tests.

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:24:25 -07:00
langsmith-forge-dev[bot]
c61d8b0376
fix: stop agent retrying commit_and_open_pr on 403 permission denied (#1123)
* fix: stop agent retrying commit_and_open_pr on 403 permission denied

Detect 403/permission-denied push failures in commit_and_open_pr and
return a PERMANENT_FAILURE message so the LLM stops retrying. Also add
prompt-level guidance to the COMMIT_PR_SECTION reinforcing this. Add
unit tests covering both the 403 and non-403 push failure paths.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: stop safety net retrying permanent push failures

Skip the after-agent PR fallback when commit_and_open_pr reports a permanent GitHub push authorization failure, while preserving fallback behavior for recoverable failures.

---------

Co-authored-by: Claude Agent <agent@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 19:37:53 +00:00
Johannes du Plessis
6ef70823be
fix(open_pr): reuse config for GitHub token; defer installation token lookup (#1227)
Calling get_github_token() without arguments always invoked LangGraph get_config internally,
which broke tests that only patch agent.middleware.open_pr.get_config and failed outside
runnable context.

Extend get_github_token with an optional runnable config mapping; the middleware passes
the config dict already resolved from get_config(). Request GitHub App installation tokens
only after detecting sandbox/repo changes worth publishing.

Fixes failing Agent unit tests in tests/test_open_pr_middleware.py.
2026-04-30 17:48:15 -07:00
langsmith-forge-dev[bot]
b5ed2a6b8b
fix: safety net middleware never fires due to key-existence check (#1051)
* fix: safety net middleware always skipped due to key-existence check

The open_pr_if_needed after-agent middleware checked `if 'success' in pr_payload`
which evaluates True for BOTH success and failure responses from commit_and_open_pr
(all responses include the 'success' key). This meant the safety net never fired.

Fix: use `pr_payload.get('success')` to check the VALUE instead of key existence.

Evidence: 6+ production traces in last 24h where commit_and_open_pr returned
success=False but the safety net silently skipped (non-fast-forward push failures,
missing GitHub token, workflow permission errors, API 500 errors).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* update

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Palash Shah <palash@langchain.dev>
2026-04-30 17:22:11 -07:00
Aran Yogesh
920a8a7624
feat: open PRs under user's name and add OpenSWE label (#1215)
* feat: open PRs under user's name and add OpenSWE label

* feat: use user token for PR authorship, add OpenSWE label, and consolidate fallback logic

* linting

* fix: address review nits for PR authorship and labeling

Fix docstring casing, add debug logging for 422 existing-PR search
fallback, tighten test type annotations, and add missing HTTPError
fallback test.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-04-29 17:34:00 -07:00
Aran Yogesh
a3a40a1bec
Revert "fix: Auto assign PRs to creator (#1211)" (#1214)
This reverts commit e9b94ac8ba.
2026-04-22 13:46:44 -07:00
Brace Sproul
e9b94ac8ba
fix: Auto assign PRs to creator (#1211)
* fix: Auto assign PRs to creator

* cr
2026-04-21 14:13:55 -07:00
Aran Yogesh
4d4f5fbfc7
fix: proxy config restored the branch yogesh/GitHub auth proxy (#1173)
* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* removing logger.info

* formatting and linting

* fix: resolve lint errors in server.py (imports, unused vars, undefined names)

* feat: use opaque proxy headers for GitHub auth in sandbox

* linting formatting and test changes

* linting

* Delete .claude directory

* Delete tests/evals directory

* fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests

* fix: restore authorship, branch_name support, and installation token for PR creation

* linitng

* fix: move installation token fetch before commit, clean up dead proxy validation code

* fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config

* fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config

* linting

* linting

* fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox
2026-04-08 15:02:52 -07:00
Aran Yogesh
9aba4d0545
Revert "feat: authenticate git operations via sandbox proxy instead of creden…" (#1170)
This reverts commit 6305e13dc6.
2026-04-07 18:45:33 -07:00
Aran Yogesh
6305e13dc6
feat: authenticate git operations via sandbox proxy instead of credential files [closes: OPE-20] (#1070)
* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* removing logger.info

* formatting and linting

* fix: resolve lint errors in server.py (imports, unused vars, undefined names)

* feat: use opaque proxy headers for GitHub auth in sandbox

* linting formatting and test changes

* linting

* Delete .claude directory

* Delete tests/evals directory

* fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests

* fix: restore authorship, branch_name support, and installation token for PR creation

* linitng

* fix: move installation token fetch before commit, clean up dead proxy validation code

* fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config

* fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config

* linting

* linting

* fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox
2026-04-07 16:41:48 -07:00
Brace Sproul
fd8e6d98ee
fix: Shell injection and ssrf issues (#1155)
* fix: Shell injection and ssrf issues

* cr
2026-04-01 12:37:35 -07:00
Aran Yogesh
86307affe4
fix: use GitHub App installation token for PR creation instead of user token (#1149)
* fix: use GitHub App installation token for PR creation instead of user token

* fix: move installation token fetch after no-changes check to avoid unnecessary API call
2026-03-30 12:47:57 -07:00
Brace Sproul
87968ab813
fix: Add scripts for getting usage, fix dual committing (#1131) 2026-03-25 13:39:33 -07:00
Aran Yogesh
150ff6f0c8
fix: allow open-swe to be triggered on any GitHuh branch (#1100)
* fix: allow open-swe to be triggered on any GitHuh branch

* formmatting
2026-03-20 10:53:33 -07:00
Brace Sproul
d7d9bc5179
fix: better custom backend support (#1071)
* fix: better custom backend support

* cr
2026-03-17 11:55:36 -07:00
agent-forge-app[bot]
d98cd9ddc1
fix: skip no_op injection when PR is committed and user is notified (#1047)
* fix: skip no_op injection when PR is committed and user is notified

- Root cause: ensure_no_empty_msg Branch 1 (empty AI message) lacked the
  same completion checks as Branch 2 (text-only AI message), causing a
  spurious no_op injection after commit_and_open_pr + user notification
- Change: add check_if_model_already_called_commit_and_open_pr AND
  check_if_model_messaged_user guard to Branch 1 of ensure_no_empty_msg
- Verified: 7/100 production traces no longer get an extra LLM call

* updated

* pass tests

---------

Co-authored-by: Auto Fix Bot <auto-fix@langchain.ai>
Co-authored-by: Palash Shah <palash@langchain.dev>
Co-authored-by: Palash Shah <35114859+Palashio@users.noreply.github.com>
2026-03-10 12:56:44 -07:00
Aran Yogesh
b735afb445
feat: add GitHub PR comment trigger and reply support [closes OPE-1] (#1014)
* feat: add GitHub PR comment trigger and reply support

* refactor: improve readability of GitHub integration

* linting

* fix: resolve github token from thread metadata and improve PR trigger flow

* fix: fall back to OAuth for GitHub webhook when no token in thread metadata

* give me commit message github integeration working without a breaking

* auth.py refactor

* fix: validate cached GitHub token before use to handle expiry

* liniting

* ci unitest formatting

* feat: post PR comments as GitHub App bot instead of user OAuth token

* resolved comments

* slack resolveed comments

* Refactor docstring and comments in get_slack_repo_config

Removed unnecessary comments and cleaned up docstring formatting.

* cr

* cr

* cr

* cr

---------

Co-authored-by: bracesproul <braceasproul@gmail.com>
2026-03-09 17:14:13 -07:00
Brace Sproul
bd52e5e09d
chore: Drop monorepo (#1029)
* chore: Drop monorepo

* cr
2026-03-06 16:10:34 -08:00