* fix: recover from mid-run sandbox death
Recreate dead sandboxes during tool execution and stop repeated unrecoverable timeout loops with a user-facing notification.
* fix: count repeated sandbox recreations
Treat consecutive sandbox recreations as an unrecovered failure streak so outages cannot loop until the model-call limit.
* feat: add optional Slack Assistants API typing status indicator
Mirrors OpenClaw's pragmatic approach: instead of rebuilding around
assistant_thread_started events, just opt into assistants.threads.setStatus
to show 'is thinking…' while the agent is working, and clear it when
post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so
it can be toggled without touching code.
* fix(slack): drop redundant clear, add status heartbeat across model calls
- Slack auto-clears the typing indicator on bot post; remove the explicit
assistants.threads.setStatus("") call from post_slack_thread_reply.
- The indicator expires after ~2 minutes; add a before_model middleware
that refreshes it on every model tick so it stays visible across long
agent runs. Reuses the existing slack_thread.{channel_id,thread_ts}
configurable already plumbed for notify_step_limit.
- chat:write is sufficient on the bot token (assistant:write is on the
way out per Slack docs); no scope or app-config change required.
* feat(slack): contextual status text + rotating loading_messages
- set_slack_assistant_status now accepts an optional loading_messages list
(capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus
payload so Slack rotates through them client-side.
- The heartbeat middleware derives a contextual status from the last
assistant message's tool calls (e.g. "searching the codebase…" after
grep, "running commands…" after execute), falling back to the default
"is thinking…" when no tool calls or unknown tool name.
- Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the
contextual status on each refresh.
* fix slack assistant status lifecycle
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: add reviewer graph + eval target wiring
- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
alongside the main `agent` graph. Reuses the same sandbox lifecycle,
GH proxy auth, and middleware primitives from `agent.server`, but with
a narrower tool set, a reviewer-specific system prompt, no
commit/push, and the `task` (subagent) tool stripped via
`_ToolExclusionMiddleware` so review stays in one context.
- New `github_comment` tool: agents call it once per issue with
`(file, line, body, severity)` and the eval scores those calls
against golden comments.
- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
*not* on the reviewer's stack — that middleware exists to enforce the
main agent's "always finalize via Slack/Linear/PR" contract, which
the reviewer doesn't have. The main agent's behavior is unchanged.
- `evals/reviewer/target.py`: send PR info as a user message, extract
every `github_comment` tool call (multiple expected per review) into
the run output.
- `evals/reviewer/judge.py`: per-example evaluator now returns a list
of metrics under `{"results": [...]}` so LangSmith averages each
numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
the UI. Dropped the broken `aggregate_pr` summary evaluator that
reached for an attribute that doesn't exist on `RunTree`.
- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
`client.list_examples(limit=N)` since `aevaluate` doesn't accept
`max_examples`.
- Makefile: `dev` and `run` targets now use `uv run` so they work
without an activated venv.
* resolve comments
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: move github workflows to gh cli
Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.
* docker ignore + snapshot and docker image updates
* updated image and instructions
* removing open_pr if needed after agent call
* fix: prevent futile retry loop when commit_and_open_pr fails with git/API errors
- Root cause: when git checkout or GitHub PR API fails, the tool returned a generic {"success": false} error with no signal that retrying is futile, causing the agent to loop 9-13+ times until hitting the 1000-step recursion limit
- Change: (1) git_checkout_branch now returns (bool, str) so the actual git error output is surfaced in the tool response; (2) checkout and PR creation failures now include "fatal": true and an explicit "Do not retry" message; (3) prompt.py COMMIT_PR_SECTION adds an explicit instruction to stop on fatal errors
- Verified: 109 unit tests pass, no regressions
* fix: skip PR safety net on fatal commit failures
* style(open_pr): ruff-format fatal retry skip condition
---------
Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
- Root cause: LLM occasionally generates strings like '1, 80' or '170, "limit": 60'
for integer fields, causing a Pydantic ValidationError and wasting an LLM turn
- Change: add SanitizeToolInputsMiddleware in agent/middleware/sanitize_tool_inputs.py
that extracts the leading integer from any string value in offset/limit before
the call reaches Pydantic validation; registered before ToolErrorMiddleware in server.py
- Verified: 14 unit tests covering all three production trace patterns pass
Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix: notify users via Slack when agent hits model call step limit
- Root cause: GraphRecursionError at 1000 steps bypassed all @after_agent
middleware including open_pr_if_needed, leaving users with no notification
- Change: Added ModelCallLimitMiddleware(run_limit=60) to intercept gracefully
before the hard recursion limit, and added notify_step_limit_reached
@after_agent middleware to post a Slack thread reply when the limit fires
- Verified: 107 existing tests pass, no regressions
* fix: harden step-limit Slack notification
Ensure the step-limit notification runs after the PR safety net and cover the new middleware behavior with focused unit tests.
---------
Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix: stop agent retrying commit_and_open_pr on 403 permission denied
Detect 403/permission-denied push failures in commit_and_open_pr and
return a PERMANENT_FAILURE message so the LLM stops retrying. Also add
prompt-level guidance to the COMMIT_PR_SECTION reinforcing this. Add
unit tests covering both the 403 and non-403 push failure paths.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: stop safety net retrying permanent push failures
Skip the after-agent PR fallback when commit_and_open_pr reports a permanent GitHub push authorization failure, while preserving fallback behavior for recoverable failures.
---------
Co-authored-by: Claude Agent <agent@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Calling get_github_token() without arguments always invoked LangGraph get_config internally,
which broke tests that only patch agent.middleware.open_pr.get_config and failed outside
runnable context.
Extend get_github_token with an optional runnable config mapping; the middleware passes
the config dict already resolved from get_config(). Request GitHub App installation tokens
only after detecting sandbox/repo changes worth publishing.
Fixes failing Agent unit tests in tests/test_open_pr_middleware.py.
* fix: safety net middleware always skipped due to key-existence check
The open_pr_if_needed after-agent middleware checked `if 'success' in pr_payload`
which evaluates True for BOTH success and failure responses from commit_and_open_pr
(all responses include the 'success' key). This meant the safety net never fired.
Fix: use `pr_payload.get('success')` to check the VALUE instead of key existence.
Evidence: 6+ production traces in last 24h where commit_and_open_pr returned
success=False but the safety net silently skipped (non-fast-forward push failures,
missing GitHub token, workflow permission errors, API 500 errors).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* update
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Palash Shah <palash@langchain.dev>
* feat: open PRs under user's name and add OpenSWE label
* feat: use user token for PR authorship, add OpenSWE label, and consolidate fallback logic
* linting
* fix: address review nits for PR authorship and labeling
Fix docstring casing, add debug logging for 422 existing-PR search
fallback, tighten test type annotations, and add missing HTTPError
fallback test.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: use GitHub App installation token for PR creation instead of user token
* fix: move installation token fetch after no-changes check to avoid unnecessary API call
* fix: skip no_op injection when PR is committed and user is notified
- Root cause: ensure_no_empty_msg Branch 1 (empty AI message) lacked the
same completion checks as Branch 2 (text-only AI message), causing a
spurious no_op injection after commit_and_open_pr + user notification
- Change: add check_if_model_already_called_commit_and_open_pr AND
check_if_model_messaged_user guard to Branch 1 of ensure_no_empty_msg
- Verified: 7/100 production traces no longer get an extra LLM call
* updated
* pass tests
---------
Co-authored-by: Auto Fix Bot <auto-fix@langchain.ai>
Co-authored-by: Palash Shah <palash@langchain.dev>
Co-authored-by: Palash Shah <35114859+Palashio@users.noreply.github.com>
* feat: add GitHub PR comment trigger and reply support
* refactor: improve readability of GitHub integration
* linting
* fix: resolve github token from thread metadata and improve PR trigger flow
* fix: fall back to OAuth for GitHub webhook when no token in thread metadata
* give me commit message github integeration working without a breaking
* auth.py refactor
* fix: validate cached GitHub token before use to handle expiry
* liniting
* ci unitest formatting
* feat: post PR comments as GitHub App bot instead of user OAuth token
* resolved comments
* slack resolveed comments
* Refactor docstring and comments in get_slack_repo_config
Removed unnecessary comments and cleaned up docstring formatting.
* cr
* cr
* cr
* cr
---------
Co-authored-by: bracesproul <braceasproul@gmail.com>