Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
(cherry picked from commit 27d90ef196)
* Adopt upstream modular webhook skeleton (#1621)
Apply the durable-interrupt-dispatch refactor: split the monolithic
webapp.py into a thin routing layer plus per-source handlers in
webhooks/{github,slack,linear}.py, and add completion.py, dispatch.py,
and reconcile.py. Reconcile fork divergence by keeping the Bedrock/
Fireworks cross-provider fallback, the no-agent-attribution prompt
policy, the dashboard-handoff re-export, and the Slack channel-info
cache. ci_autofix is restored on the new dispatch model in a later
commit.
Refs: #80
* Port fork webhook security delta onto modular handlers
Re-apply the fork's security customizations that #1621 did not carry:
Linear webhook replay protection (freshness window on the signed
webhookTimestamp), per-repo token-cache binding threaded through the
thread token resolvers, the INTERNAL_BOT_LOGINS self-check in the
review-finding-reply path, and a user-mapping cache refresh before
email resolution on the issue and PR-comment paths (multi-replica
staleness). Existing fork security tests pass unchanged.
Refs: #80
* Restore CI auto-fix on the modular dispatch model
Bring back ci_autofix.py and the ci_monitor graph that #1621 deleted,
re-wiring the fork's security-reviewed PR-babysitting onto the new
structure: the CI-event, autofix-toggle, and review-feedback handlers
move into webhooks/github.py and the github_webhook router re-gains the
check_run/check_suite/workflow_run/status routing plus the autofix
command and actionable-review branches. Auto-fix runs now dispatch
through dispatch_agent_run (durability + completion webhook) while
keeping the deliberate batch-while-busy skip-rule via
get_thread_active_status. Restore langgraph.json's ci_monitor entry and
the fork autofix tests (dispatch mock + import paths re-pointed).
Refs: #80
* Reformat and update docs for the modular webhook split
Point CLAUDE.md and deploy/MIGRATION.md at the new webhooks/ modules
and the dispatch/completion/reconcile contract, and mark the
user-mapping cache-refresh fix as applied on the GitHub handlers.
Refs: #80
* Restore reject backstop for autofix dispatch
A burst of near-simultaneous CI events for one head SHA can slip past
the busy-check before the dedupe SHA is recorded, so dispatch the
autofix path with multitask_strategy=reject (dev's prior platform
default) to drop duplicate concurrent creates instead of letting them
interrupt each other. Also make the completion failure-reply dedup
claim-then-post and drop the unreachable interrupted branch.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
* fix: enforce a replay window on Linear webhooks (AUTHZ-001)
verify_linear_signature accepted any correctly-signed body with no freshness
check, so a captured request could be replayed indefinitely. Parse the
signed webhookTimestamp (Unix ms) and reject requests outside a 60s window,
failing closed when the field is missing or malformed — mirroring the Slack
verifier.
* fix: stop leaking upstream auth-error bodies into user comments
get_github_token_for_user folded the raw upstream response text into the
error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the
full body server-side only and return a generic "GitHub auth failed (status
<code>)". Also document the accepted shared-installation-token blast radius on
the bot-token-only path (AUTHZ-003).
* fix: bind sandbox and token caches to repo to prevent thread-id collision
A PR head-branch name is attacker-controllable and get_thread_id_from_branch
derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01).
The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on
thread_id alone, and a cached sandbox was reused after only an echo-ping, so a
different repo's webhook could bind to another thread's sandbox or token.
Without changing the persistent thread-id scheme:
- Persist the bound repo (owner/name) in thread metadata on sandbox creation and
refuse to reuse a sandbox whose bound repo does not match the current event
(SandboxRepoMismatchError); the in-memory proxy also carries the binding.
- Bind the GitHub-token cache entries to their repo and evict on a cross-repo
read so a colliding thread_id cannot be served another repo's token.
- Thread repo through the reviewer and the webhook token resolvers.
* fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04)
The instance role and the GitHub deploy app role granted s3:ListBucket on the
whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts)
only ever lists under releases/, so add a StringLike s3:prefix=releases/*
condition. GetBucketLocation has no s3:prefix in its request context, so it
moves to its own unconditioned statement. Also document the accepted F-2
cross-env existence-oracle residual on BatchGetSecretValue.
* chore: suppress test-fixture credential false positive; document AUTHZ-002
Add a machine-level suppression for the fake Datadog key in the
test_team_credentials encryption-roundtrip fixture (CWE-798, not a real
credential). Clarify that the within-org thread-write path is intentional by
design (AUTHZ-002) — comment only, no behavior change.
* fix: casefold repo-binding keys to avoid spurious cross-repo mismatch
GitHub owner/name are case-insensitive. Casefold the owner/name key on both the
write (binding) and read (compare) sides — repo_cache_key and the metadata
bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate
same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2).
* fix: stop leaking upstream auth body in unexpected-result branch
The 2xx-but-missing-token/url branch echoed the parsed upstream response body
into the user-facing error. Return a generic message and log response_data
server-side only, mirroring the existing HTTPStatusError fix (Gap 4).
* fix: fail closed for unbound-legacy sandboxes and catch repo mismatch
Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no
recorded bound_repo (a pre-binding legacy thread, post-deploy) previously
reconnected-and-served the sandbox to the current repo, then rebound it. Now
fail closed: drop the stale id and recreate a fresh sandbox bound to this repo,
logging a reconnect-with-missing-binding event. A sandbox is never served to a
repo unless its binding is known and matches; new threads bind on first run
unchanged.
Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints,
log it for alarming, and surface a clean sanitized error instead of letting an
opaque deep-stack exception crash-loop the worker.
* chore: suppress test-fixture credential false positive in token-TTL tests
Add a machine-level suppression for the fake "ghp_secret" GitHub token used by
the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and
not a valid PAT; scoped to the unit test only.
* feat(reviewer): surface dashboard UI link on PR reviews [INF-0000]
Post a transient "review in progress" comment (with an "Open in Web"
dashboard link) when a reviewer run starts, then delete it once the
review lands. The published review body now carries the same
"Open in Web" link, so the link persists on the review itself.
The transient comment's id is tracked in reviewer thread metadata
(status_comment_id) so it can be deleted on completion.
* refactor(reviewer): inline dashboard URL helper, drop redundant future import [INF-0000]
Reviewer/push runs execute in a worker process, but the GitHub token was
cached by the webhook handler in the API server process — a different
process — so the worker's in-process cache was always cold. resolve_github_token
then failed (User not authenticated / Unknown source: github_push) and only the
app-token fallback kept reviews working, noisily.
The reviewer always acts as the GitHub App (open-swe[bot]), so resolve the
installation token directly at run start, scoped to the repo. This also bypasses
org SAML enforcement that blocks user OAuth tokens. Drop the now-dead
cross-process cache writes in the webhook reviewer-dispatch handlers, and stop
leave_failure_comment raising on the github_push source.
Open SWE Review now runs from automated PR triggers, so drop the old Slack/GitHub review keyword entrypoints and keep PR comments on the regular agent path.
* feat: add Slack Open in Web link
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: skip web link for Slack reviewer runs
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: accept include_dashboard_link kwarg in Slack reviewer test double
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
* fix: resolve reviewer head_sha from thread metadata, not frozen run config
A push that lands while a reviewer run is in flight is delivered as a
queued message into that run. The run's configurable is frozen at
creation, so its head_sha still names the commit the run was created for
— not the commit just pushed. publish_review then anchored the GitHub
review to the stale commit and regressed last_reviewed_sha to it, and
add_finding/update_finding stamped findings with the stale SHA.
Persist the current head in thread metadata at every reviewer dispatch
(both the ready-for-review and push paths, before they branch to create
a run or queue a message), and add resolve_review_head_sha() which
prefers the metadata head over the run config. Wire it into
publish_review (review commit_id + last_reviewed_sha), add_finding
(first_seen_sha) and update_finding (last_confirmed_sha). Falls back to
the run config when metadata carries no head (first review, eval, tests).
* fix: persist head_sha in manual review dispatch (trigger_pr_review_from_ref)
resolve_review_head_sha prefers metadata[head_sha] over the run config,
and the push/ready dispatchers write it — but trigger_pr_review_from_ref
(Slack/GitHub @open-swe review, request_pr_review tool) created a run
with a freshly-fetched config head while leaving metadata's head stale
from a prior dispatch. A manual re-review at a newer commit would then
resolve to the old head and publish/advance findings against it.
Persist head_sha in that dispatch's metadata write too, so every
run-creating reviewer dispatch keeps metadata in sync with the head its
run targets. Caught by the Open SWE reviewer on this PR.
* Replace hardcoded GitHub-email map with Store-backed user mapping
Move the static GITHUB_USER_EMAIL_MAP to a Store-backed bidirectional
mapping (GitHub login <-> work email <-> optional Slack ID) with an
in-process cache, self-service onboarding, and admin management.
- agent/dashboard/user_mappings.py: Store CRUD + login/email/slack-id
indexes, sync cache readers for hot paths, async fallthrough, and a
bulk_import that preserves existing richer records.
- Migrate all read sites (auth.py, agent_overrides.py, authorship.py,
github_comments.py, webapp.py x2) off the dict.
- Unmapped Slack tags now run on the GitHub App installation token
(use_installation_token_fallback) and get an ephemeral "link your
GitHub account" prompt carrying the Slack id + email via a signed
account-link token threaded through the OAuth state.
- OAuth callback completes a self-service (org-gated) mapping from that
token, falling back to the verified GitHub email.
- Admin CRUD endpoints + one-time legacy import; dashboard UI section.
- Legacy dict retained only as the import payload (no longer read).
Tests: mapping store, account-link round-trip + completion, mapped vs
unmapped Slack flows; existing trust-gate tests updated to prime cache.
* Address review: cold-cache email resolution + stale alias de-indexing
- agent_overrides: add resolve_login_from_email_async that falls through to
the Store on a cold cache; use it at the async repo-resolution call sites
(Slack repo config, Linear comment, owner-metadata) so a mapped user still
resolves to their GitHub login + dashboard default_repo on a fresh worker.
- user_mappings.upsert_mapping: de-index the existing login before re-indexing
so a changed email/Slack id no longer leaves stale aliases resolving to the
login in-process.
- Tests for both fixes; update Slack repo-config test to patch the async resolver.
* feat: reconcile reviewer findings with PR threads
* fix: harden reviewer finding reply handling
* fix: queue reviewer finding reply body
Ensure review-comment replies that arrive during an active reviewer run include the sanitized reply body in the queued reassessment prompt.
* fix: apply reviewer reply formatting
Apply the repository formatter so the reviewer reply handling fix passes CI format checks.
* fix: route PR review requests through agent tool
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: align public repo gate test with agent routing
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: allow app-token PR review requests
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat(dashboard): restructure Open SWE Review tab + wire create_prs
Restructures the dashboard around two related changes the reviewer settings
have been asking for:
- Wire profile.create_prs. Defaults to true (opt-out); when off the system
prompt gets a `Pull Request Policy Override` section telling the agent
to push the branch and notify with the branch URL instead of opening a
PR. Removes the noop Slack Notifications / Allow Artifacts / First Name
/ Last Name controls and their schema fields.
- Repositories opt-in for Open SWE Review. New per-team enabled list
stored in the LangGraph Store (`["enabled_review_repos"]`). Every
reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review`
which AND-combines the existing env allowlist with the dashboard list.
Default is empty (opt-in) — admins enable repos per-installation from
the new Repositories page nested under Open SWE Review.
- Open SWE Review tab now mirrors the Cursor "rules" pattern: main page
shows installation rows + a Rules entry; both drill into nested pages
(/review/repositories/$owner and /review/styles) with a back link.
- Adds the new logo/favicon assets shipped from sidebar + html head.
Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults
`is_review_repo_enabled` to True for existing allowlist tests.
* fix(dashboard): make main content scroll independently of the sidebar
Outer flex container was min-h-svh, so it grew with main's content and the
whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden
so the sidebar stays put and only <main> scrolls.
* fix(dashboard): make disabled repo toggles obviously disabled
Switch's disabled state used opacity-50 against a muted background, so
the not-admin state looked nearly identical to the off state. Bump to
opacity-40 + grayscale, and wrap each repo toggle in a span carrying a
native hover tooltip explaining why it's disabled.
* fix(switch): handle base-ui's data-disabled state
base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute)
when disabled, so Tailwind's disabled: variant never matches and the
button keeps its cursor-pointer + clickable look. Mirror the styling
under the data-[disabled] variant and add pointer-events-none so the
disabled state is both visible and actually unclickable.
* feat(dashboard): paginate per-installation repository list
20 repos per page with Prev / page X of Y / Next controls at the bottom.
Pager only renders when there are more than 20 repos. Page resets to 0
when navigating between installations.
* feat(dashboard): global default model selectors for Agent + Reviewer
Adds team-wide default model + reasoning effort for both agents in the
Admin tab so operators can switch models without redeploying.
Resolution chain:
Agent: hardcoded -> LLM_MODEL_ID env -> team default -> user profile
Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable
Team defaults live in team_settings and are validated against the
SUPPORTED_MODELS allowlist + the model's supported reasoning efforts.
'Inherit from env' clears the override and falls back to LLM_MODEL_ID.
* refactor(models): drop LLM_MODEL_ID env in favour of the team default
The team default is now the single source of truth for the runtime model
choice; per-user (agent) and per-call configurable (reviewer) selections
still win on top. When no admin has touched the team default, it surfaces
the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the
admin UI's dropdown is always pre-populated with a sensible value.
The Admin UI loses the 'Inherit from env' option since there is no longer
an env layer to inherit from.
* chore(models): set hardcoded fallback to gpt-5.5 medium
Decouple the team-default boot value (gpt-5.5 / medium) from each model's
ProfileForm-suggested default_effort so we can change one without nudging
the other. The Opus xhigh default for new user profiles is unchanged.
* feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings
- Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new
description copy that matches the screenshot. Legacy stored values
fall back to 'every_push' on read so the UI never shows an unknown
selection.
- Add a 'Coming soon' badge + greyed-out + disabled state on the
controls that don't have runtime consumers yet: Trigger Mode,
Autofix Mode, Autofix Severity Threshold, and Automatically fix CI
failures. SettingsRow grew a comingSoon prop to keep this consistent.
- My Settings drops the noop PR Preferences section and adds a Sign
Out button. preferred_pr_destination is removed from the profile
schema; old records get the field popped on next write.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Avoid treating regex-parsed Slack text as authoritative routing or allowlist input so repository mentions remain model context and access is bounded by installation permissions.
* feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints
Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering:
- GitHub App OAuth login → JWT cookie session (cross-domain ready)
- profile CRUD against LangGraph Store with model+effort validation
- admin gate via CONFIGURED_ADMINS
- /repos via /user/installations using the user's encrypted OAuth token
CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the
Vercel-hosted frontend can call the LangSmith deployment with credentials.
* feat: apply dashboard profile model/effort overrides in get_agent
Look up the triggering user's GitHub login from config (direct field or
GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store,
and apply default_model + reasoning_effort to make_model when both are
valid. Effort 'max' is captured on the profile but not yet wired through —
the OpenAI Reasoning Literal doesn't accept it.
* feat: ui/ TanStack Start dashboard for profile config
Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template,
base-ui primitives, Tailwind v4). Three routes:
- /login — Sign in with GitHub (links to /dashboard/api/auth/login)
- /profile — Edit default model, reasoning effort, default repo
- /admin — Admin-only: list users and edit other profiles
API client (src/lib/api.ts) uses credentials: include so the osw_session
cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL
points at the LangSmith deployment.
Effort options re-render when the model changes; 'max' on Opus 4.7 is
captured on the profile but ignored downstream until anthropic reasoning
is wired through make_model.
* feat: searchable Combobox for default repo picker
Replaces the Select with a base-ui Combobox so users can filter by typing,
the popup is wider than the trigger so full owner/repo names are readable,
and the list caps at max-h-80 to stay on screen.
* fix: address review comments + wire default_repo and Anthropic thinking
Security/correctness fixes from PR review:
* Open redirect: validate `redirect_to` in `/auth/login` against
`DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it
into the state JWT. Anything off-allowlist falls back to the dashboard
base URL. (PR #1302 r3250054386)
* Login CSRF: bind the OAuth `state` to the requesting browser. At
`/auth/login` we generate a fresh nonce, set it as a short-lived
HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and
embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback`
we require the cookie nonce to hash-match the state JWT's nonce_hash
(constant-time compare). (PR #1302 r3250054395)
* RMW race in profile vs token writes: split storage into two
namespaces — `["profiles"]` for user-editable settings and
`["oauth_tokens"]` for the encrypted GitHub token. Each upsert now
only writes its own namespace so an in-flight profile save can no
longer clobber a fresh token from a concurrent re-login (and vice
versa). (PR #1302 r3250054393)
* /repos pagination: follow `Link: rel="next"` for both
`/user/installations` and per-installation `/repositories` with
per_page=100, capped at 1000 items. (PR #1302 r3250054401)
Feature wires:
* default_repo: applied as a fallback in `get_slack_repo_config` (after
explicit-repo / thread metadata, before the env defaults) and in the
Linear webhook (after comment-body extraction, before team mapping).
Both paths resolve the triggering user's GitHub login via
GITHUB_USER_EMAIL_MAP and read the profile's default_repo.
* Anthropic "thinking" effort: `make_model` now accepts a `thinking`
kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max}
to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is
anthropic. OpenAI path still ignores "max" since the Literal doesn't
accept it.
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]
Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* webapp: forward installation-token expiry to reviewer cache writes
The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Now that assistant.threads.setStatus carries the "is thinking…"
indicator, the per-run greeting phrase + auto-unfurl on the LangSmith
link are just visual noise stacking on top of it. The trace reply now
posts only `<url|View trace>` + a tip, with link unfurling off so the
smith.langchain.com card no longer appears.
Adds an `unfurl_links`/`unfurl_media` knob to post_slack_thread_reply_with_ts
(default-on to preserve behaviour for every other caller).
* feat: add optional Slack Assistants API typing status indicator
Mirrors OpenClaw's pragmatic approach: instead of rebuilding around
assistant_thread_started events, just opt into assistants.threads.setStatus
to show 'is thinking…' while the agent is working, and clear it when
post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so
it can be toggled without touching code.
* fix(slack): drop redundant clear, add status heartbeat across model calls
- Slack auto-clears the typing indicator on bot post; remove the explicit
assistants.threads.setStatus("") call from post_slack_thread_reply.
- The indicator expires after ~2 minutes; add a before_model middleware
that refreshes it on every model tick so it stays visible across long
agent runs. Reuses the existing slack_thread.{channel_id,thread_ts}
configurable already plumbed for notify_step_limit.
- chat:write is sufficient on the bot token (assistant:write is on the
way out per Slack docs); no scope or app-config change required.
* feat(slack): contextual status text + rotating loading_messages
- set_slack_assistant_status now accepts an optional loading_messages list
(capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus
payload so Slack rotates through them client-side.
- The heartbeat middleware derives a contextual status from the last
assistant message's tool calls (e.g. "searching the codebase…" after
grep, "running commands…" after execute), falling back to the default
"is thinking…" when no tool calls or unknown tool name.
- Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the
contextual status on each refresh.
* fix slack assistant status lifecycle
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat(open-swe): trigger reviewer agent from `@open-swe review` PR comment
Mirrors the Slack `@open-swe review` flow on GitHub: a comment containing
`@open-swe review` (optionally followed by a PR URL) on a PR triggers the
reviewer agent. Without a URL it reviews the commenting PR; with a URL it
targets that PR. Works for `issue_comment`, `pull_request_review_comment`,
and `pull_request_review` events, gated by the existing reviewer repo
allowlist and reusing `trigger_pr_review_from_ref`.
* fix(open-swe): require URL after `@open-swe review`, don't swallow trailing text
The previous regex matched any non-whitespace token after `review`, including
across newlines. Comments like `@open-swe review\nthanks!` parsed as
`(True, "thanks!")`, which then failed PR-URL parsing and was silently
dropped — the user got no review and the comment never reached the regular
PR-comment handler.
Restrict the optional URL token to `https?://\S+` so non-URL trailing text
falls through to `process_github_pr_comment` instead of being eaten by the
review-command branch. Adds regression tests for the multiline and
trailing-word cases.
When a Slack user kicks off a PR review with `@open-swe review <pr-url>`,
the reviewer agent now posts a one-line summary back to the Slack thread
when it finishes — either "No issues found" or "found N potential
issue(s)" with a link to the GitHub review.
The reviewer agent has no Slack tools by design, so the summary is sent
host-side from `publish_review` after the GitHub review POST succeeds.
The Slack channel/thread_ts is persisted on reviewer thread metadata at
trigger time and read back on publish. Re-reviews triggered by push
events stay silent in Slack to avoid noise on the original thread.
* feat: implement reviewer findings, publish_review, and watch mode
Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:
- Findings as first-class state on the reviewer thread metadata
(`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
ranges, suggestion text for ```suggestion blocks, github_review_comment_id
for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
compute_diff_line_set for in-diff validation, extract_diff_hunk for
caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
ranges fail at creation, not at GitHub-publish), `update_finding`,
`list_findings`, `publish_review`. The reviewer agent's tool list is
swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
one POST /reviews call with body + inline comments + ```suggestion blocks,
per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
before the agent's first model call (warm- and cold-path symmetric);
computed diff and in-diff line set passed via runnable config; system
prompt rewritten for the single-evolving-findings model, severity ladder,
in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
added to supported events. New `process_github_push_event` resolves the
open PR for the pushed branch, gates on the reviewer thread's `watch`
flag, builds a re-review configurable, and triggers a run on the same
canonical thread. `process_github_pr_close` toggles watch on
closed/reopened. `set_reviewer_thread_metadata` is called on first
review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
legacy {file, line, body, severity} shape the judge expects) and passes
the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
publish rendering + GraphQL resolve, and watch-mode webhook handlers
(push triggers re-review only when watching, idempotent on unchanged
head SHA, PR close disables watch). Updated existing reviewer-webhook
tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
into REVIEWER_DESIGN.md.
* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL
Address PR #1253 review findings:
- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
(`option no-prefix takes no value` — every prep run was failing
silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
now uses three-dot `base...head` (the merge-base diff GitHub renders
on Files-changed) so we don't pick up changes that landed on the base
branch after the PR diverged. Re-review delta keeps two-dot
`last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
`github_review_comment_id`. Without this, watched re-reviews
re-posted every previously surfaced finding, and only the most-recent
duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
`/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
is the canonical endpoint; the old form 404s, so comment ids were
never stored and watch-mode resolution couldn't run.
Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.
* fix(reviewer): default publish cap from 15 to 4
A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
* Use gh for reviewer inline comments
* Trigger reviewer on PR review requests
* Format reviewer webhook test
* Add GitHub repo allowlist
* Separate reviewer repo allowlist
* only implement allowlist for reviewer
* feat: move github workflows to gh cli
Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.
* docker ignore + snapshot and docker image updates
* updated image and instructions
* removing open_pr if needed after agent call
* feat: add GitHub PR comment trigger and reply support
* refactor: improve readability of GitHub integration
* linting
* fix: resolve github token from thread metadata and improve PR trigger flow
* fix: fall back to OAuth for GitHub webhook when no token in thread metadata
* give me commit message github integeration working without a breaking
* auth.py refactor
* fix: validate cached GitHub token before use to handle expiry
* liniting
* ci unitest formatting
* feat: post PR comments as GitHub App bot instead of user OAuth token
* resolved comments
* slack resolveed comments
* Refactor docstring and comments in get_slack_repo_config
Removed unnecessary comments and cleaned up docstring formatting.
* cr
* cr
* cr
* cr
---------
Co-authored-by: bracesproul <braceasproul@gmail.com>